Skip to content

anthropic

Claude's hillclimb cut support costs to about a fifth

Promtime

On 14 support tickets it never saw during the search, a setup tuned by Claude scored 90.5% against the original 78.6%, at about one fifth of the cost. The run is described in Anthropic's post on the Claude blog, which introduces two commands in the claude-api skill for Claude Code: /claude-api build-eval and /claude-api hillclimb.

At a glance

  • The claude-api skill in Claude Code gains two commands: build-eval interviews you and builds an evaluation inside your codebase, and hillclimb then improves your application against it, one patch at a time.
  • In Anthropic's support-ticket example, the climb moved from Opus 4.8 at high effort and 4.6 cents per ticket to Sonnet 5 on low effort at about 1 cent.
  • Both worked examples are Anthropic's own internal runs, the held-out set is 14 tickets, and the post prints no confidence intervals for its own headline figures.

If you have not been following along, the claude-api skill gives Claude Code guidance on Anthropic's APIs and general tips for working with Claude. According to Flowtivity, it is an open-source install distributed from Anthropic's GitHub skills repository. Explainx says it had already covered /claude-api hillclimb as a cost-search tool in an earlier piece, "Claude Platform cost and effort", before this companion post on eval design.

A good eval keeps the strongest model well below 100%

Anthropic lists four traits of a trustworthy eval. Tasks mirror production, not whatever is easy to generate or grade. Stronger models and more effort score higher, and if they don't, the usual suspects are ambiguous tasks or a miscalibrated grader. The best model at the highest effort sits well below 100%. Run-to-run variance is low, and that includes variance hiding in config or in leftover state from an earlier trial, such as a file or a git history.

A task that fails every run, however many replicates you do, is a sign that it is impossible or ambiguous. The post also warns against adversarial sampling. If you pick cases because today's model fails them, you end up measuring that model's failure fingerprint. You should be able to say why a task is hard before you include it. User traffic alone may also skew easy, because users try what they expect to work.

Build-eval starts with production transcripts and only then synthesizes cases

When you run /claude-api build-eval, Claude interviews you and pauses for your approval at set points. It gathers inputs in a fixed order. First come production transcripts, after questions about retention and sensitive data, then bug reports and support tickets, then five to ten cases you write by hand, and finally cases synthesized from your codebase. Claude builds a simple page that lists every input and waits for you to confirm them.

Next it proposes the cheapest grader that fits. Constrained outputs get a code check, such as an exact match or JSON validated against a schema. Open-ended outputs get an LLM judge that scores against a rubric written as checkable claims, not a 1-to-5 scale. You choose the judge model, and it should not be the model you are testing.

Claude grades a handful of cases and asks whether you would have scored any of them differently. During the baseline runs it grades the same output twice to see whether the verdict changes. It also checks for timeouts, API errors and cut-off answers. If the baseline already scores about 95% or higher, it warns you.

Opus 4.8 at 4.6 cents per ticket became Sonnet 5 at about 1 cent

The cost example ran on an internal customer support benchmark of 44 tickets, with 30 used for the search and 14 held out. It started on Opus 4.8 at default (high) effort, with 74.4% decision accuracy on the search tickets at 4.6 cents per ticket. The climb began by auditing the prompt and removing mandatory tool-call rituals, a scratchpad step and some contradictory rules.

It then tried Opus 5.5 on low effort, which reached 87.8% at 1.9 cents per ticket. Part of that saving is pricing. On Opus 5.5, input and output tokens cost 20% less than on Opus 4.8, and cache reads cost 60% less. Dropping a tier, Sonnet 5 on low effort scored 88.9% at about 1 cent per ticket.

After Claude added routing rules and a refund-cap cross-reference to the prompt, Sonnet 5 reached 98.9% on the search tickets at about the same cost. On the 14 held-out tickets, the final configuration scored 90.5% against the original setup's 78.6%.

The claude-api skill's own score rose from 66% to about 88%

The second example tests the claude-api skill itself, on an eval built from Anthropic's documentation. It started at 66%. With access to the docs and SDKs, the hillclimber found eight features the skill did not cover, and adding sections for them lifted the score to 74%. Fixing errors in the C# and Java type tables took it to 77%.

After two stalled rounds, Claude sorted the remaining failures by cause. The content was already in the skill, but Claude kept writing older API shapes from its trained priors. The fix was a table near the top of the skill that maps old forms to current ones. One example is fixed-budget extended thinking, which the API now rejects on recent Opus models, to adaptive thinking. The score reached 80%.

The same pass caught flawed tasks. In one, the task asked for code that catches one error type, while its grader wanted a chain of at least three, so Claude reworded the task. In another, the grader contradicted the docs, and tests against the real API showed the docs were right. According to Explainx, the climb to about 88% took 24 rounds.

A hillclimb patch survives only if the held-out set improves too

First, Claude asks what to optimize: performance, or cost while performance holds. It also asks what it may change. The options are the system prompt, skills or instruction files, tool descriptions, model and effort settings, and harness code. It then splits the eval at random into train and test sets and checks that the eval's noise is smaller than the smallest improvement you would act on.

In each round, Claude reads the train transcripts and proposes one patch that targets the root cause of a failure. If train improves while test stays flat, it treats that as overfitting and reverts. It also reverts regressions. It works like a student who memorized last year's exam: only the score on a fresh paper counts.

Failure content never gets pasted into the prompt. At the end, Claude leaves your code at the version that did best on the test set. It reports the gain with confidence intervals, and if the gain is within noise, it recommends against merging.

The caveat is the evidence. Both examples are Anthropic's internal runs, and the cost case holds out just 14 tickets. Oddly, a post that insists on confidence intervals prints none for its own headline figures. And the gap between 98.9% on search tickets and 90.5% on held-out ones is likely the kind of spread its own overfitting check is there to catch.

Results on evals Anthropic didn't build

Both commands are available now in Claude Code through the claude-api skill. The post gives no results from teams outside Anthropic. So the open question is whether a jump like 78.6% to 90.5% holds up on evals the company did not design. The report's own rule is the one to watch: a gain within noise does not get merged.

Related stories

  1. Claude Code grades your plugin against a no-plugin run
  2. Claude Code skips AGENTS.md when telemetry is off
  3. Claude Mods ship in weeks, and you can flip them on now
  4. /resume carries CLI sessions into the desktop app
  5. /design shows up in Claude Code as research preview
  6. Output tokens cost roughly 5x input in Claude Code

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.