Skip to content

anthropic

Sonnet 5.5 hits third place by writing 410M tokens

Claude News

In the Artificial Analysis Intelligence Index, Claude Sonnet 5.5 at max effort generated 410M output tokens. That is more than four and a half times the 88M used by the median model in its class. The extra thinking earned a score of 56, well above the class median of 26, and third place out of 216 models.

At a glance

  • Anthropic released Claude Sonnet 5.5 on September 28, 2026. Artificial Analysis tested its max-effort reasoning configuration, which accepts text and image input and has a 1M-token context window.
  • List prices are $2.00 per 1M input tokens and $10.00 per 1M output tokens, which are exactly the class medians. Artificial Analysis puts the blended rate at $1.54 for a 7:2:1 mix of cache hits, input and output.
  • The catch is volume. At max effort, one Intelligence Index task costs $7.60 on average. Artificial Analysis lists output speed as N/A, so nobody knows yet how long all those tokens take to generate.

If you have not been following, Sonnet 5.5 arrives after a busy few months. According to ScriptByAI, Claude Sonnet 5 came out on June 30, 2026, and Claude Opus 5.5 followed on September 22, 2026. Anthropic has also built a Mythos-class tier above Opus. It started with Claude Fable 5 and Mythos 5 on June 9, 2026, followed by Fable 5.1 and the restricted-access Mythos 5.1 on September 1.

Sonnet 5.5 scores 56 on an index where its class median is 26

The Artificial Analysis Intelligence Index v4.3.2 combines 10 evaluations into one number: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1. Several of them are agentic, covering knowledge work, real-world work tasks, SaaS workflows and coding in a terminal.

On that mix, Sonnet 5.5 scores 56 in its adaptive reasoning, max effort configuration. The class median is 26, and Sonnet 5.5 ranks third of 216 models. The page is labelled Adaptive Reasoning, Max Effort, Default Fallback. Artificial Analysis notes that a non-reasoning variant may also exist.

The 216 does not include every model on the site. Proprietary models are compared with proprietary and open-weights models in the same price band, based on a blended 3:1 input-to-output price. At $2.00 and $10.00, Sonnet 5.5 falls into the top band, above $1 per 1M tokens.

The output bill: 410M tokens and $7.60 per task

The per-token prices look ordinary. On Anthropic's API, Sonnet 5.5 costs $2.00 per 1M input tokens and $10.00 per 1M output tokens, and Artificial Analysis places both at the class median. Cached prompts get a 90% discount. At a 7:2:1 mix of cache hits, input and output, the blended rate is $1.54 per 1M tokens.

The difference is in volume. While running the Index, Sonnet 5.5 at max effort generated 410M output tokens against a class median of 88M, and Artificial Analysis calls that very verbose. The average cost per task is $7.60. That figure is weighted across the 10 evaluations and counts input, cache hits, cache writes, reasoning tokens and answer tokens.

For comparison within Anthropic's lineup, ScriptByAI lists Opus 5.5 at $4 per million input tokens and $20 per million output tokens. That is double Sonnet 5.5's list price on both lines.

Where do 410M tokens come from?

They come from reasoning as well as from answers, and Artificial Analysis splits its token count into those two categories. According to AI Arte, reasoning tokens are billed as output tokens and count toward max_tokens. The billing is the same even when the thinking text is never returned to you. In the same source's description, Claude restates the question, tries several approaches, checks intermediate results and drops paths that do not hold up.

The control for how much thinking happens has changed over time. According to hidekazu-konishi.com, extended thinking first shipped in Claude 3.7 Sonnet in February 2025. The Claude 4.x generation, starting in May 2025, added an effort parameter for trading speed against capability. AI Arte describes the newer adaptive thinking, where the model itself decides whether to think and how deeply instead of working to a fixed budget_tokens value.

Think of a taxi meter that keeps running while the driver considers alternative routes in his head. You pay for every route he considered, not only the one he drove. Max effort is the setting where Sonnet 5.5 is allowed to consider the most routes, and the meter runs the whole time.

The page leaves two gaps. Speed is listed as N/A, so the page does not tell you how long 410M tokens take to arrive, and its text does not say how much of the $7.60 goes to reasoning and how much to answers. In our view, the headline price understates what this configuration really costs. The per-token rate sits at the median, but the token count is more than four times the median, and the bill follows the count.

The speed number still missing

The next figure to watch is output speed, which Artificial Analysis currently lists as N/A for this configuration. Once it is measured, you can weigh 410M tokens in time as well as in money. Sonnet 5.5 is available through 3 API providers, and Artificial Analysis notes that pricing can vary by provider, so the $2.00 and $10.00 list prices may not hold everywhere. No date for the speed measurement has been given.

Related stories

  1. Opus 5.5 leads the index and burns 260M tokens doing it
  2. Max effort adds nothing to Fable 5.1's ARC-AGI-2 score
  3. Opus 5 tops Fable 5 on OSWorld 2.0 at a third of the price
  4. Claude Fable 5: access and capabilities
  5. Effort levels recalibrated in Opus 4.8
  6. Leaks put Claude Sonnet 5.5 at $2 input, release next week

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.