Skip to content

anthropic

Claude Sonnet 5.5 keeps Sonnet 5's price, cuts cost per task

Promtime

Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, an agentic coding test. That puts it above Opus 5.5 at 66.4% and far above Sonnet 5 at 10.3%. The price did not change, but according to Anthropic, the new model uses so many fewer tokens that a typical task costs up to 30% less.

At a glance

  • Claude Sonnet 5.5 is the second model in the Claude 5.5 family after Opus 5.5. It is aimed at well-scoped everyday tasks, bug fixing, documents, slides and spreadsheets, and Haiku 5.5 is due in the coming weeks.
  • List prices stay at $2 per million input tokens and $10 per million output, half of Opus 5.5's $4 and $20. Anthropic says Sonnet 5.5 generates output 30%+ faster than Sonnet 5.
  • The catch: it is the first Sonnet with cyber safeguards, so higher-risk security tasks visibly fall back to Sonnet 5. Anthropic also still rates Opus 5.5 clearly stronger at open-ended work.

If you haven't been following: according to The News Stack, Opus 5.5 launched last week, and Sonnet 5.5 followed on Monday as the second model in the family. Anthropic describes Sonnet as the faster, cheaper complement to Opus. Opus is for complex work that needs careful judgment, Sonnet for well-scoped tasks, and Haiku for high-volume, cost-sensitive applications. Haiku 5.5 is the only one still to come.

Sonnet 5.5 scores 1,844 on GDPval-AA, two points behind Opus 5.5

GDPval-AA, from Artificial Analysis, ranks models by Elo score on real-world tasks across 44 occupations and nine major industries. Sonnet 5.5 scores 1,844. Opus 5.5 scores 1,846, Sonnet 5 scores 1,449 and OpenAI's GPT-6 Sol scores 1,487. AA-Briefcase gives the same order: 1,811 for Sonnet 5.5, 1,822 for Opus 5.5, 1,359 for Sonnet 5 and 1,483 for GPT-6 Sol.

The gap to Opus is small on other tests too. On OSWorld 2.1, a computer-use test, Sonnet 5.5 scores 80.1% and Opus 5.5 scores 81.8%, while Sonnet 5 scored 57.0%. On Humanity's Last Exam with tools, Sonnet 5.5 gets 64.5%, Opus 5.5 67.7% and Sonnet 5 54.9%. On Chartography, a chart-reading test run without tools, Sonnet 5.5 gets 61.6%, Opus 5.5 64.4%, Sonnet 5 15.6% and GPT-6 Sol 53.6%.

Anthropic also says it is the first Sonnet model to beat Pokémon Red working only from screenshots. In one internal test, Anthropic gave it a public company's quarterly earnings materials, call transcripts and a slide template, and asked for a 10-slide operating review. Two experts judged the first draft ready to send as is.

On Terminal-Bench, the coding score jumps from 10.3% to 70.6%

Terminal-Bench 4.0 measures complex, multi-step professional tasks in a command-line interface. CursorBench 4.0 uses tasks from real Cursor coding sessions. There, Sonnet 5.5 scores 55.5%, Opus 5.5 57.8% and Sonnet 5 34.1%. On FrontierCode, Anthropic says Sonnet 5.5 at High effort scores 10 points higher than Sonnet 5 at the same setting, for about one fifteenth of the cost per task.

The News Stack describes FrontierCode, from Cognition, as a test of whether a code change could be merged without human edits. It reports Sonnet 5.5 at 52.1% at its second-highest effort setting, against 54.4% for Opus 5.5 and 49.3% for GPT-6 Sol. Early testers noticed that Sonnet 5.5 grouped more tool calls together than Sonnet 5 did, so it needed fewer steps. Daniel Vogel, COO of Epic Games, said the model held up on a system design audit and a data flow review:

The new model managed tens of thousands of lines of code for gameplay system architecture, kept responses snappy, handled multi-hour tasks, and delivered with less prescriptive prompting.

The list price stays at $2 and $10 per million tokens

Sonnet 5.5 keeps Sonnet 5's list price: $2 per million input tokens, $10 per million output tokens and $0.20 per million for cache reads. Cache writes cost $2.50. Opus 5.5 charges $4 for input, $20 for output and $5 for cache writes, and the same $0.20 for cache reads.

According to The News Stack, Anthropic cut Opus 5.5's price to those $4 and $20 rates, and OpenAI halved prices for GPT-6 Sol and Luna on the same day. The outlet notes that Sonnet 5.5's list price now matches GPT-6 Sol, OpenAI's second-best model after Astra.

Curtis Allen, a principal engineer at Slack, said that without any prompt changes, Sonnet 5.5 beat Sonnet 5 on almost all of Slack's offline Slackbot evals. It did so in fewer steps and with about 14% fewer output tokens.

How the same price becomes a cheaper task

What you pay for a task is the price per token times the number of tokens the model uses, and the price has not changed. According to Anthropic, the token count has: Sonnet 5.5 typically needs far fewer tokens for the same work. Think of a taxi with the same meter rate and a driver who knows a shorter route.

The effort setting is the second lever. At lower effort, Claude answers faster and uses fewer tokens. At higher effort, it reasons for longer and checks its work more thoroughly, so cost and score usually rise together. Claude Code and the Claude apps default to Medium, and the Claude Platform defaults to High.

Anthropic says that on several benchmarks, Sonnet 5.5 at Low or Medium effort beats Sonnet 5's best score for about a tenth of the cost per task. On Terminal-Bench at Medium, it far exceeds Sonnet 5's best score for less than a tenth of the cost. At higher settings, the company says, it can perform comparably to Opus 5.5 at a similar cost.

Higher-risk security requests fall back to Sonnet 5

Anthropic says Sonnet 5.5's cybersecurity capabilities are comparable to Opus 5's, so it ships with safeguards similar to those on Opus 5.5. Routine bug fixing is unaffected, but higher-risk cyber tasks visibly fall back to Sonnet 5. The biology safeguards are the same as Sonnet 5's, and some microbiology and virology requests may be flagged in error. Organizations can apply to the Life Sciences Verification Program, and an expanded Cyber Verification Program is coming soon.

Sonnet 5.5 is also the first Sonnet model with classifiers that stop attackers from extracting its reasoning. The target is distillation, where attackers use thousands of fake accounts to copy a model's capabilities. An expanded feature called preserved thinking ties Claude's thinking to the account that created it, which changes what happens when you switch accounts mid-session in Claude Code. On an automated behavioral audit of roughly 1,850 scenarios, Sonnet 5.5 matches or beats Sonnet 5 on most measures, though Opus 5.5 does slightly better overall.

Most of these numbers come from Anthropic's own testing, and the 30% saving is an "up to" figure. The News Stack adds a caveat on Terminal-Bench: the benchmark's maintainers noted that Sonnet 5 sometimes ran into timeouts and token limits, which helps explain its low score. In our view, the default setting matters as much as the model: the Claude Platform starts at High effort, so API users who never change it may not see the cheapest version of these results.

Before you switch to claude-sonnet-5-5

The model is live now on all platforms, including Amazon Web Services, Google Cloud and Microsoft Azure, with zero data retention. If you run Sonnet with thinking off, you will need to switch to the new between_tools setting before migrating. The News Stack notes that Opus 5.5 already rejects requests that turn thinking off entirely. Haiku 5.5 is promised "in the coming weeks," and the expanded Cyber Verification Program "soon." Neither has a date.

Related stories

  1. 90.0% on ARC-AGI-2 at $4.49 per task
  2. CodeRabbit: Opus 5.5 trades 9 missed bugs for 11 new ones
  3. Opus 5.5 lands 40% cheaper to run than Opus 5
  4. Xiaomi's new MiMo models carry an unverified top-6 claim
  5. xAI's new transcriber marks speakers and drops the ums
  6. Qwen3.8-Omni-Flash takes video and audio in a 1M window

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.