Best practices

The Builder Cannot Grade Its Own Homework

Most agent runs fail in a boring way. The model produces something plausible, explains why the choices were reasonable, and stops. You asked for great. You received a defense of pretty good. The missing piece is not more adjectives in the prompt. It is a second pair of eyes that never saw how the work was made.

What Matt Shumer actually changed

In July 2026 Matt Shumer posted a Three.js first-person shooter that Claude Code built from a short prompt. The run lasted many hours, wrote roughly 55,000 lines of code, and generated textures, meshes, animation, and sound in code. The clip reached 3.8 million views. Skeptics reran the prompt. When they kept his structure, they got working games. When they kept only the enthusiasm, they got demos.

He later named the structure the Gauntlet Loop. The prompt itself is short on purpose. It gives a destination and a bar. It tells the agent to split the work, fan out subagents, and keep a harsh critic comparing the real output with a real reference. It does not include an architecture. That is the point. A detailed plan replaces the model’s judgment with yours, and then you wonder why the run cannot surprise you.

The rule that travels

The builder has seen every compromise. It remembers why the headline is timid, why the lighting is flat, why the paragraph hedges. That memory makes it a gifted explainer of its own work. It is a poor judge of the work. A critic with a clean context window, the bar, and the rendered artifact has none of those memories. It can only look.

That rule is bigger than games. A landing page can be compared with a live page in the same category. A research brief can be compared with a paper whose structure you trust. A backend change can be compared with a test suite or a latency number. The bar has to be inspectable. “Make it premium” is not a bar. It is a mood, and moods pass everything.

What we add when we teach it as a skill

Matt Shumer’s method is the core: goal, bar, split, independent critic, no fixed round count. Packaging it as a skill adds two gates that prevent the failure we see most often. An interview that refuses to start without a specific reference. A preflight that confirms the critic can actually see (screenshot, render, test output). Those gates do not tell the agent how to build. They stop it from looping against nothing.

A skill also gives you a place to write the brake. This loop is expensive. The original run was many hours on a strong model with high effort. You are the person who stops it. Put the ceiling in the first answer of the interview so the agent treats it as a checkpoint, not a dare.

When not to use it

If the task is cheap, skip the gauntlet. Formatting a commit, renaming a file, and drafting a reply do not need a critic fleet. Use this when the artifact is what people will judge in public, and when you can name something that already does that job well. One piece is enough for a first run. Extra pieces multiply cost.

Stay connected

Never miss a post

Updates on format changes, community features, and skill building.