Skip to content

anthropic

Delegation made Opus 5 refuse tasks it solves on its own

Promtime

A four-model Claude stack running inside Claude Code scored 78% on Terminal-Bench 2.1, seventh place, and spent $1,178 doing it, roughly twice what the leader spent. The benchmark hands the agent 89 terminal tasks with five attempts each.

Three security tasks never got solved once. Opus 5, working as the executor, refused to take them on across all 14 attempts. The same Opus 5 in a plain session with no orchestration solved all three, six for six. Count those and the run lands at 80.5% and third place.

The verifier agent only ran when the orchestrator decided to call it, which happened in 76% of runs. With verification, tasks were solved 91% of the time; without it, 43%.

The Opus 5 executor burned 58% of the budget, $689. Runs with 3 to 5 handoffs between agents solved 90% of tasks at $2.41 per attempt. At six handoffs or more, the solve rate dropped to 50% and the cost rose to $9.51.

Related stories

  1. Claude Code sessions keep running after you close the laptop
  2. Claude Code skips AGENTS.md when telemetry is off
  3. Anthropic tells Opus 5.5 users to delete "think hard" lines
  4. CodeRabbit: Opus 5.5 trades 9 missed bugs for 11 new ones
  5. Claude Code hands paid users a limit reset through Oct 22
  6. Claude Code's weekly cap drops 17% on September 14

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.