Skip to content

anthropic

Claude Sonnet 5.5 took #3 with 410M tokens of output

Promtime

According to Artificial Analysis, Claude Sonnet 5.5 at max effort produced 410M output tokens on the way to a score of 56 on the Intelligence Index. For comparable models, the medians are 88M tokens and a score of 26. That score puts Anthropic's model, released on September 28, 2026, third of the 216 models in its comparison class.

At a glance

  • Artificial Analysis profiles Claude Sonnet 5.5 in its Adaptive Reasoning, Max Effort, Default Fallback configuration, a proprietary reasoning model with text and image input, text output and a 1M-token context window.
  • List prices are $2 per 1M input tokens and $10 per 1M output tokens, exactly the class medians, and one Intelligence Index task at max effort costs $7.60 on average.
  • Output speed was not measured and the field reads N/A. Artificial Analysis itself calls the 410M-token run very verbose, so we still have no figure for how long this configuration takes.

In case you haven't been following: Sonnet is the middle line of Anthropic's lineup, between the smaller Haiku and the larger Opus. Artificial Analysis tests the models it tracks against the same set of evaluations and gives each configuration its own page. The page for this one carries a long name, Adaptive Reasoning, Max Effort, Default Fallback, and says a non-reasoning variant may also exist.

A score of 56 puts Sonnet 5.5 third of 216 models in its class

Version 4.3.2 of the Artificial Analysis Intelligence Index combines 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1. Several of them are agentic, covering knowledge work, real-world work tasks, SaaS workflows and coding in a terminal. So the index measures a lot more than chat quality.

At max effort, Sonnet 5.5 scores 56 on this index. The median for comparable models is 26, and Artificial Analysis calls the result well above average. The model ranks #3 of 216 in its class. One detail: CritPt, the physics reasoning evaluation, is currently marked as under review on the page.

Sonnet 5.5 wrote 410M output tokens where the tier median is 88M

During the Intelligence Index run, Sonnet 5.5 at max effort generated 410M output tokens. The median among reasoning models in a similar price tier is 88M, which makes this configuration more than four and a half times as talkative. Artificial Analysis describes it as very verbose, at the higher end of its tier.

Output tokens here cover two things: the final answer and the reasoning the model writes before it. The token charts on the page show these as separate parts. Speed was not measured, and the output tokens per second field reads N/A. That leaves the page with no figure for how long a max-effort answer takes to arrive.

Prices sit at the class medians of $2 and $10 per million tokens

On Anthropic's API, Sonnet 5.5 costs $2.00 per 1M input tokens and $10.00 per 1M output tokens, and both prices match the class medians exactly. Cached prompts get a 90% discount. At a blended 7:2:1 ratio of cache hits, input and output, Artificial Analysis puts the effective rate at $1.54 per 1M tokens.

At max effort, one Intelligence Index task costs $7.60 on average, which puts the model #98 of 216 on cost. It is available through 3 API providers and accepts up to 1M tokens of context, about 1,500 A4 pages in 12-point Arial. Anthropic has not disclosed the model's size or parameter count.

A $7.60 task bill is built from five kinds of tokens

The per-task figure is not a list price. Artificial Analysis takes the tokens a model used in each evaluation and splits them into input, cache hits, cache writes, reasoning and answer tokens. It prices each kind, divides the total by the number of tasks and weights the result by that evaluation's share of the index.

Picture two taxis with the same meter. The rate per kilometre is identical, but the driver who takes the long way hands you a bigger fare. Sonnet 5.5 charges the median rate per token, and at max effort it writes far more tokens than its median peer.

The comparison class is also set by price. Proprietary models are compared with proprietary and open-weights models in the same price band, using a blended 3:1 input-to-output price. There are three bands: below $0.15, from $0.15 to $1, and above $1 per 1M tokens. Reasoning models like this one are compared against both reasoning and non-reasoning models.

The page leaves some gaps. It has no speed figure and does not explain what Default Fallback means in the configuration name. It also lists a verbosity rank of #102 of 216 without reconciling that with its own very-verbose label. In our view, the token appetite matters more than the list price: the per-token rates are ordinary, the token count is what drives the per-task bill, and it is the one part you can't negotiate.

Waiting on a max-effort speed number

Artificial Analysis gives no date for filling in the output speed field. Until it does, nobody can say how long a 410M-token run takes in practice. Provider pricing is the second number to watch: 3 API providers serve the model, and the page notes that prices may differ between them. The non-reasoning variant the page mentions has no scores here yet.

Related stories

  1. Blocked Claude requests cost money again, in three areas
  2. CodeRabbit: Opus 5.5 trades 9 missed bugs for 11 new ones
  3. Claude Code hands paid users a limit reset through Oct 22
  4. Opus 5.5 lands 40% cheaper to run than Opus 5
  5. Claude Code's weekly cap drops 17% on September 14
  6. First OpenAI win on Vending-Bench, at almost 3x the cash

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.