Claude Code writes the eval, grader included

Claude Code can now write the test that grades your own app: a runnable eval with a set of inputs, a way to run the app on each one, and a grader that scores the output. Anthropic's Claude developer team announced it on X, and according to the GitHub listing, the command stops for your sign-off before it spends any money.
At a glance
- The new /claude-api build-eval command in Claude Code produces a runnable eval, and a companion /claude-api hillclimb command, according to GitHub, then works on improving the app itself.
- Hillclimbing works like tuning a recipe by taste: you change one thing, rerun the eval, keep whatever raised the score, and repeat the loop on your application.
- According to the GitHub listing, the build-eval command pauses for sign-off on inputs, grading method and cost before the first paid run, though the announcement itself gives no typical price for a run.
If you have not been following, Anthropic's engineering blog has published a run of pieces on evaluation methodology. According to that blog, "Demystifying evals for AI agents" came out on January 9, 2026, "Designing AI-resistant technical evaluations" on January 21, 2026 and "Quantifying infrastructure noise in agentic coding evals" on February 5, 2026, alongside a piece on eval awareness in Claude Opus 4.6's BrowseComp performance.
/claude-api build-eval returns inputs, a runner and a grader
The build-eval command ends with three pieces: a set of inputs, a way to run your app on each input, and a grader that scores the results. According to the GitHub listing, it gets there through a four-step interview, and before the first paid run it stops so you can sign off on the inputs, the grading method and the cost.
Per the same listing, Step 0 pins down what is being evaluated, Step 1 decides where prompts come from (an existing eval, transcripts or synthesized cases), Step 2 sets the grading method, and Step 3 produces a runnable script with measured cost. Anthropic's blog on claude.dev adds that Claude samples inputs in a set order: production transcripts first, after asking about retention, then other sources.
What makes an eval worth climbing?
According to Anthropic's blog on claude.dev, a well designed eval has four traits. Its tasks mirror production, scores improve with stronger models and more thinking, and the most capable model at the highest effort still lands well below 100%, which the post calls "passable" headroom. The fourth trait is low run-to-run variance.
The same claude.dev post traces high variance to ambiguous tasks or to a grader that gives different verdicts on identical output. Variance can also hide in configuration, such as effort not applied consistently, or in leftover environment state, like a file or a git history that hands the agent the answer.
The claude.dev post also warns about adversarial sampling. Model capability is jagged, so picking cases only because today's model fails them risks measuring that model's failure fingerprint instead of what is actually hard or valuable. Its rule is simple: include a task only if a human can say why it is hard.
How does hillclimb change the app?
It runs a loop, according to the GitHub listing. The /claude-api hillclimb command reads the eval's failures, applies a change, and reruns the eval on train, validation and test splits each round, stopping at a budget and stopping condition you approve. In recipe terms, the test split is the guest who tastes only the final dish.
The listing also describes a third command, /claude-api cost-optimize, which ranks ways to cut API spend without lowering quality and, on request, applies a change and measures it against your eval. Both build-eval and hillclimb load eval-audit.md, a health checklist covering task design, harness design, metrics hygiene, grader design, and whether the eval can detect the change being tested.
Per GitHub, the guides live in the skill's shared/evals/ directory: build-eval.md, eval-audit.md, eval-hillclimb.md and cost-hillclimb.md, plus runner-scaffold.mjs and build-report-lite.mjs, which generate a static report.html. The CLI extracts these files to a temporary path each session.
The inputs leave out two things: the size of the bill and the size of the gain. Neither the announcement nor the listing gives a typical cost per run or a single before-and-after score from hillclimb, so you learn both only on your own app. In our view, the sign-off step is the right design for a tool that spends your API money in a loop, because the price shows up before the first run starts.
Update to 2.1.280 first. According to the GitHub listing, the commands need Claude Code CLI version 2.1.280 or later. No sample budgets or typical score gains have been given, so the first real numbers will come from teams running build-eval on their own apps. The question to watch is whether test-split scores hold up after many hillclimb rounds, or whether gains made on the train split stop carrying over.
Related stories
- Without CLAUDE.md, Claude Code 2.1.277 reads AGENTS.md
- Claude Code tags gateway requests by class and agent
- Claude Code 2.1.269 grades plugins with a scored eval run
- Claude Code 2.1.268 makes /cost match gateway rates
- Pop-out panes surface in the Claude Code desktop app
- Claude Code 2.1.267 caps effort level across providers
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
