Skip to content

model-releases

Opus 5.5 leads the index and burns 260M tokens doing it

Claude News

Artificial Analysis spent $8,708.20 running its full test suite on Claude Opus 5.5 at max effort, and the model wrote 260M output tokens along the way. According to Artificial Analysis, that bill bought first place out of 206 models on its Intelligence Index, with a score of 58 against a median of 25.

At a glance

  • Anthropic released Claude Opus 5.5 on September 22, and the adaptive reasoning version at max effort takes text and images, outputs text and holds a 1M-token context window.
  • Pricing is $4.00 per 1M input tokens and $20.00 per 1M output tokens, twice the class medians of $2.00 and $10.00, with cost per index task at $5.98.
  • The listing shows no output speed figure for this configuration and one API provider, and Anthropic has not disclosed the model's size or parameter count for the proprietary weights.

If you have not been following, the Opus line has moved quickly this year. According to TechCrunch, Opus 5.5 arrived just two months after Opus 5 shipped on July 24. Anthropic said Opus 5 set a new state of the art on coding and knowledge-work evaluations such as Frontier-Bench and GDPval-AA, and gave greatly improved performance for the same cost as Opus 4.8.

Opus 5.5 scores 58 on an index where the class median is 25

The Artificial Analysis Intelligence Index v4.3.2 combines 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1. Opus 5.5 at max effort scores 58 on it, first of 206 models in its comparison class.

According to the Artificial Analysis methodology, the index is a weighted average over four categories: Agents at 30%, Coding at 20%, Scientific Reasoning at 20% and General at 30%. The same methodology estimates the 95% confidence interval at less than ±1%, based on experiments with more than 10 repeats on certain models.

The comparison class matters too. Proprietary models are compared with proprietary and open weights models in the same price range, using a blended 3:1 input/output price, with bands below $0.15, from $0.15 to $1 and above $1 per 1M tokens. Reasoning models like this one are compared across both reasoning and non-reasoning models.

Getting to 58 took 260M output tokens and $8,708.20

Across the index run, Opus 5.5 at max effort generated 260M output tokens. The median for reasoning models in its price tier is 92M, so the model wrote nearly three times as much, which Artificial Analysis describes as very verbose.

The full index run cost $8,708.20, and the weighted average cost per Intelligence Index task came to $5.98. Artificial Analysis builds that figure from input, cache hit, cache write, reasoning and answer token prices, divided by task count and weighted by each evaluation's index weight. The page's own summary calls Opus 5.5 amongst the leading models in intelligence, but somewhat expensive compared with models of similar price.

Output costs $20 per million tokens, double the class median

Through Anthropic's API, Opus 5.5 costs $4.00 per 1M input tokens and $20.00 per 1M output tokens, against class medians of $2.00 and $10.00. Cached prompts get a 95% discount. At a blended 7:2:1 ratio of cache hit, input and output tokens, the price works out to $2.94 per 1M tokens.

Measured against its own family, the picture shifts. According to TechCrunch, the predecessor charged $25 per million output tokens, other metrics showed similar price drops, and the new model is faster to run, reflecting an overall drop in the compute needed to serve it. Artificial Analysis lists one API provider for this model.

What does max effort actually change?

It changes how much work the model does before it answers. According to Anthropic, Claude models expose an effort parameter that trades intelligence against speed and token cost, and the charts for Opus releases show performance changing with that setting.

The dial is not new. According to a timeline on GitHub, the 4.x generation, which began with Claude Opus 4 and Claude Sonnet 4 in May 2025, introduced both the effort parameter and the 1-million-token beta context window.

Think of it as telling a contractor how many times to measure before cutting: more measuring usually means a better cut and a bigger bill. The listing name also includes "adaptive reasoning" and "default fallback", and the page defines neither term.

The gaps matter for anyone planning a budget. With no output speed figure, you cannot tell how long a 260M-token workload would take, and the page does not explain what the fallback does or when it kicks in. In our view, this configuration reads as a ceiling test more than an everyday default for Claude Code or API work, given nearly three times the median token count at double the median output price.

Waiting on Sonnet and Haiku 5.5

According to TechCrunch, Anthropic said Sonnet 5.5 and Haiku 5.5 would follow in the coming weeks with similar performance improvements. No specific date has been given, and no pricing for either model has been named. Once they ship, the open question is how their index scores and token counts compare with the 58 and 260M that Opus 5.5 posted at max effort.

Related stories

  1. Opus 5.5 aced a test suite at 3.4x GPT-6 Sol's cost
  2. Opus 5.5 takes the default seat in Claude Code
  3. Opus 5.5 matches Fable 5.1 on most work for less money
  4. Fable 5.1 tops Artificial Analysis at 66, costs 20% more
  5. Max effort adds nothing to Fable 5.1's ARC-AGI-2 score
  6. Cache reads get 75% cheaper on Claude Fable 5.1

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.