Skip to content

anthropic

Sonnet 5.5 edges past Opus 5.5 on Terminal-Bench 4.0

Claude News

On Terminal-Bench 4.0, Sonnet 5 solved 10.3% of the tasks, while its successor Claude Sonnet 5.5 solves 70.6%, ahead of Opus 5.5's 66.4% on the same agentic coding test, according to Anthropic's announcement. The price per token has not changed, but Anthropic says the new model costs up to 30% less per task.

At a glance

  • Anthropic has released Sonnet 5.5 as the second model in the Claude 5.5 family, available today on the Claude Platform, Claude Code, AWS, Google Cloud and Microsoft Azure, with Haiku 5.5 due in coming weeks.
  • Pricing stays at $2 per million input tokens and $10 per million output, but the model uses fewer tokens per task and generates output more than 30% faster than Sonnet 5.
  • Anthropic concedes that Opus 5.5 is still clearly stronger at complex, open-ended work, and under new safeguards higher-risk cybersecurity requests on Sonnet 5.5 visibly fall back to Sonnet 5.

If you haven't been following: according to Explainx, Sonnet 5 launched on June 30, 2026 as claude-sonnet-5 and became the default on the Free and Pro plans, with introductory pricing of $2/MTok input and $10/MTok output through August 31, 2026. Opus 5.5 opened the Claude 5.5 family, and Anthropic said it costs 40% less to run than Opus 5. Anthropic describes Sonnet 5.5 as the faster, lower-cost complement to Opus 5.5.

Sonnet 5.5 lands two points below Opus 5.5 on GDPval-AA

GDPval-AA v2.1 tests real-world tasks across 44 occupations and nine industries. On it, Sonnet 5.5 scores 1844. Opus 5.5 scores 1846, GPT-6 Sol 1487 and Sonnet 5 1449. AA-Briefcase v1.1 keeps the same order: 1822 for Opus 5.5, 1811 for Sonnet 5.5, 1483 for GPT-6 Sol and 1359 for Sonnet 5.

On CursorBench 4.0, which uses tasks from real Cursor sessions, Sonnet 5.5 scores 55.5%. Sonnet 5 scores 34.1% and Opus 5.5 scores 57.8%. On OSWorld 2.1 computer use it posts 80.1%, against 57.0% and 81.8%. With tools on Humanity's Last Exam, the three score 64.5%, 54.9% and 67.7%.

On chart recognition without tools, Sonnet 5.5 goes from Sonnet 5's 15.6% to 61.6%. Opus 5.5 scores 64.4% and GPT-6 Sol 53.6%. Sonnet 5.5 is also the first Sonnet to beat Pokémon Red working only from screenshots. In one internal test, two experts judged its first draft of a 10-slide operating review, built from a public company's earnings materials, call transcripts and a slide template, ready to send.

On FrontierCode at High effort, Sonnet 5.5 gains 10 points at a fifteenth of the cost

Terminal-Bench 4.0 asks a model to finish complex, multi-step tasks in a command line. At Medium effort, which is the default in the Claude apps, Sonnet 5.5 far exceeds Sonnet 5's best score for less than a tenth of the cost per task. On FrontierCode at High effort, it scores 10 points higher than Sonnet 5 at the same setting, for about one fifteenth of the cost per task.

Early testers said Sonnet 5.5 picks up a codebase quickly. In head-to-head runs it batched tool calls together more often than Sonnet 5, so it took fewer steps and cost less. Daniel Vogel, chief operating officer of Epic Games, said the model held up on a system design audit and a data flow review:

The new model managed tens of thousands of lines of code for gameplay system architecture, kept responses snappy, handled multi-hour tasks, and delivered with less prescriptive prompting.

Sonnet 5.5 costs half of Opus 5.5 per input and output token

Input, output and cache-read prices match Sonnet 5 at $2, $10 and $0.20 per million tokens, and cache writes cost $2.50. On Opus 5.5, input costs $4, output $20 and cache writes $5, while cache reads cost the same $0.20.

Your bill is the number of tokens multiplied by the price per token. Sonnet 5.5 usually needs far fewer tokens for the same work, partly because it batches tool calls into fewer steps. Think of a taxi that charges the same rate per kilometre but whose driver knows a shorter route.

Curtis Allen, a principal engineer at Slack, said that without any prompt changes Sonnet 5.5 beat Sonnet 5 on almost all of Slack's offline Slackbot evals. It did so in fewer steps and with about 14% fewer output tokens.

Effort is the second dial. At lower settings Claude answers faster and uses fewer tokens, while at higher settings it reasons for longer and checks its work more thoroughly. Claude Code and the apps default to Medium, the Claude Platform to High. Anthropic says Sonnet 5.5 complements Opus 5.5 best at lower effort.

Higher-risk cyber requests on Sonnet 5.5 fall back to Sonnet 5

Anthropic says Sonnet 5.5's cyber capabilities are comparable to Opus 5's, which makes it the first Sonnet to ship with cyber safeguards and fallbacks. You can still find and fix bugs in routine development, but higher-risk tasks visibly fall back to Sonnet 5. Biology safeguards match Sonnet 5's, and some microbiology and virology requests may be flagged in error.

On Anthropic's automated behavioral audit of roughly 1,850 scenarios, Sonnet 5.5 improves on or matches Sonnet 5 on most measures of alignment, misuse resistance and honesty. It is the least likely of Anthropic's models to probe the limits of its containers, though Opus 5.5 does slightly better on the audit overall.

New classifiers block reasoning extraction, a defense against distillation attacks run through thousands of fake accounts. Thinking stays with the organization that generated it, so if you switch accounts mid-session in Claude Code, Claude rereads the session and generates new thinking. Zero data retention remains available, as it was for Sonnet 5.

The 10.3% baseline needs context: according to Explainx, Sonnet 5 scored 80.4% on Terminal-Bench 2.1 at launch, and a Tech Blog write-up describes 4.0 as a harder 66-task set that knocks frontier models back to roughly 60%. In our view, Anthropic's own caveat that Opus 5.5 stays stronger at open-ended work is a better guide than the Terminal-Bench lead. Note too that "up to 30%" is a ceiling, and Anthropic does not give the typical saving.

Haiku 5.5 has no date yet

Haiku 5.5, built for high-volume and cost-sensitive work, is due in the coming weeks, so the 5.5 family photo still has an empty chair. The expanded Cyber Verification Program is only promised for "soon". Pro, Max and Team users can apply last Tuesday's reset any time before Oct 22. If you run Sonnet with thinking off, switch to the between_tools setting before moving to claude-sonnet-5-5.

Related stories

  1. Leaks put Claude Sonnet 5.5 at $2 input, release next week
  2. Max effort adds nothing to Fable 5.1's ARC-AGI-2 score
  3. Opus 5 tops Fable 5 on OSWorld 2.0 at a third of the price
  4. Claude Fable 5: access and capabilities
  5. Effort levels recalibrated in Opus 4.8
  6. Opus 5.5 aced a test suite at 3.4x GPT-6 Sol's cost

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.