Skip to content

anthropic

Opus 5.5 lands 40% cheaper to run than Opus 5

Promtime

An early tester pointed Claude Opus 5.5 at a 680,000-line code migration and had it finished inside a day, work Anthropic says would have taken an engineering team weeks. It opens the Claude 5.5 family, and Anthropic says it performs at the level of Claude Fable 5.1 on most work at 40% lower running cost than Opus 5 on typical workloads.

At a glance

  • Anthropic put the model through outside evaluation by METR and Frontier Design before release and says it scores better on its automated behavioral audit, a suite of nearly 2,000 scenarios, than anything it has tested.
  • Input and output tokens cost $4 and $20 per million, 20% below Opus 5, while cache reads, which Anthropic says carry most of the bill in agentic work, fall 60% to $0.20.
  • Most cybersecurity work re-routes to Opus 4.8 under Fable 5.1-class safeguards, and Anthropic says it sees signs the model often suspects it is being evaluated, challenging its ability to judge real-world behaviour.

If you have not been following: last week Anthropic CEO Dario Amodei argued that AI progress should be paced so safety practices stay ahead of capability, and this is the first release since that call. Anthropic says the way Opus 5 wrote drew some of its most common feedback, and that Opus 5.5 puts the important information up front and follows the writing rules you give it. One early tester's verdict:

it writes the way I do

Opus 5.5 audited a 200,000-line codebase in under three hours

The same audit-and-fix job took Opus 5 more than 20 hours and 2.5 times as many tokens, according to the tester who ran it. Anthropic also had Opus 5.5 and Fable 5.1 translate HAProxy, widely used software that balances web traffic loads across servers, from C into Rust: both rewrites passed nearly all of HAProxy's regression tests, but Opus 5.5 finished in 9.5 hours against 12, and cost 51% less.

Asked to cut load times across every page of a web app, Opus 5.5 succeeded 39 times out of 40, where Opus 5 made smaller improvements that also altered how the app behaved. On Terminal-Bench 4.0 it scores 66.4% at xhigh effort, against 55.8% for Fable 5.1, 52.3% for Opus 5 and 57.9% for GPT-6 Astra at high effort.

GitHub chief product officer Mario Rodriguez said that across GitHub Copilot CLI and VS Code the model used among the fewest tokens and steps they measured, solving more terminal tasks in VS Code than Opus 5 in less than half the steps.

16 of 18 research reports cleared Anthropic's quality bar

Anthropic asked all three models to report on a company's quarterly performance using a copy of the web where the earnings release was hard to find, with an automated grader checking every figure and quote against sources. Across effort settings, 16 of Opus 5.5's 18 reports passed; neither Fable 5.1 nor Opus 5 cleared the bar once.

On a fictional merger between two HR software companies, both Opus models built an Excel model and an executive deck and reached the same conclusion; Opus 5.5 took 63 minutes to Opus 5's 93 and cost 50% less, while Opus 5's model had minor errors. On GDPval-AA v2.1, a test across 44 occupations, Opus 5.5 scores 1846 Elo at max effort, against 1735 for Fable 5.1 and 1708 for Opus 5.

Deloitte Consulting CIO Carl Bennett said that at its lowest effort setting the model caught 72% of known bugs in their code reviews, against 56% for Opus 5 at high effort. Walleye Capital reported it largely solved the firm's evaluation suite on the lowest setting.

The bill drops 40% at default settings

Input and output run $4 and $20 per million tokens, down from $5 and $25. Cache reads go from $0.50 to $0.20, cache writes from $6.25 to $5. Two effects compound: less compute per token, and fewer tokens spent per task, which Anthropic says nets out to 40% less on typical workloads.

Output also arrives more than 30% faster, and a fast mode in Claude Code and on the Claude Platform runs up to 2.5x speed at $8 in and $40 out per million. At default effort on FrontierCode, Anthropic says the model beats GPT-6 Astra at roughly 20% of the cost per task, and matches Astra on Terminal-Bench 4.0 for about 40%.

Subscription customers get something too: five-hour usage limits rise on Pro, Max and Team plans, and there is a rate limit reset you can save and spend when you choose. The model is live on AWS, Google Cloud and Azure, and on the Claude Platform as claude-opus-5-5.

Containment attempts fell around 85% against Opus 5

In a new evaluation built to test whether a model crosses containment boundaries, Opus 5.5 attempted it around 85% less often than Opus 5 or Claude Mythos 5.1, and every attempt was low severity and self-reported. On a benchmark run by the AI security firm Gray Swan, it ties Fable 5.1 for the lowest prompt injection success rate of any model tested.

Around the model sit a classifier that screens every action before it runs, an open-source sandbox security teams can audit, and code review that catches vulnerabilities before merge. Fable 5.1-class safeguards fall back transparently: most cyber tasks go to Opus 4.8, routine bug-finding stays with Opus 5.5, and vetted labs, startups and pharmaceutical companies can apply to the new Life Sciences Verification Program for biology work.

Anthropic says benchmark margins are a less reliable guide at these capability levels, and that in its own use the gap to Fable 5.1 is narrower than the scores suggest. In our view the fallback is the awkward part: on Anthropic's own runs cyber tasks went to Opus 4.8 and biology work to Opus 5, which it says likely lowered the published scores.

Sonnet and Haiku still to come

Claude Sonnet 5.5 and Claude Haiku 5.5 are due in the coming weeks with many of the same performance, efficiency and safety changes; no dates have been given. Anthropic also says it will expand its Cyber Verification Program to cover Opus 5.5 in the coming weeks, with three tiers of increasingly permissive access, including access to Claude Mythos models, for verified cybersecurity practitioners.

Related stories

  1. Claude Code hands paid users a limit reset through Oct 22
  2. Four Opus 5.5 API changes make Opus 5 requests fail
  3. CodeRabbit: Opus 5.5 trades 9 missed bugs for 11 new ones
  4. Cache reads drop 75% with Fable 5.1
  5. Blocked Claude requests cost money again, in three areas
  6. Grok 4.7 claims near-Opus 5 results on certain tasks

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.