Skip to content

benchmarks

A four-model Claude ensemble scored 78% on Terminal-Bench 2.1, at double the cost of first place

Claude News

The run placed seventh with a bill of $1,178, roughly twice what the top entry spent. The division of labor: Fable 5 plans and delegates, Haiku 4.5 scouts, Opus 5 edits code, Sonnet 5 verifies the work.

Three security tasks went unsolved because the executor refused to run them: zero out of 14 delegated attempts. In a control run without delegation, Opus 5 solved the same tasks six times out of six. Credit those three and the ensemble lands at 80.5% and third place.

Verification did most of the heavy lifting. Tasks that went through the verifier passed 91% of the time, versus 43% without the check. Opus 5 in the executor role accounted for 58% of total spend, or $689.

First place on the board still belongs to Claude Code with Fable 5 at 83.8%.

Related stories

  1. Claude Code matches Codex on SWE-Bench Pro at twice the cost
  2. Auditing one Claude session: 7.5M tokens, 4,038 tool calls
  3. Claude Code adopters merged 24% more PRs, study finds
  4. Claude Code sends 33k tokens before it even reads your prompt; OpenCode uses 7k
  5. Claude went from bug hunt to abuse report on New Year's Eve
  6. Claude Code Projects threads can now run on your machine

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.