Skip to content

anthropic

Haiku 5.5 ranks #2 of 182 and burns 440M tokens doing it

Claude News

Claude Haiku 5.5 at max effort streams 243.4 tokens per second, yet in Artificial Analysis testing its first token took 415.30 seconds to arrive, close to seven minutes. The same run ranked the new Haiku second of 182 models in its class, with an Intelligence Index score of 43.

At a glance

  • Artificial Analysis benchmarked the reasoning version of Claude Haiku 5.5, released October 7, 2026, at max effort through Anthropic's API, comparing it with proprietary and open-weights models in its price range.
  • Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens, below tier medians of $0.25 and $0.90, and an average index task cost $0.21.
  • The catch is volume: Haiku 5.5 produced 440M output tokens across the index against a 100M median, and as a reasoning model it bills its thinking tokens as output.

If you have not been following, the Claude 5.5 family opened with Opus 5.5 on September 22, 2026, and Sonnet 5.5 joined on September 28, according to ClaudeKit. The Next AI Model tracker reports that on September 22 Anthropic promised Sonnet 5.5 and Haiku 5.5 "in the coming weeks", and counts a roughly 357-day gap since Haiku 4.5, against a usual 291 days for the line.

Anthropic's own announcement dates Haiku 4.5 to Oct 15, 2025, when, by Anthropic's account, it matched Sonnet 4 on coding, computer use and agent tasks and scored 73.3% on SWE-bench Verified. Anthropic now calls Haiku 5.5 a significant step up over Haiku 4.5 across coding, tool use, computer use and agents.

Haiku 5.5 at max effort scores 43 where the tier median is 13

The Artificial Analysis Intelligence Index v4.3.2 bundles 10 evaluations, among them Terminal-Bench 4.0, SciCode, Humanity's Last Exam, CritPt, GDPval-AA v2.1 and the long-context AA-LCR v1.1. Haiku 5.5 at max effort scored 43 on it, second of 182 models in its comparison class, where the median score is 13.

The comparison class needs a word of explanation. Artificial Analysis ranks proprietary models against proprietary and open-weights models in the same price band, using a blended 3:1 input-to-output price, and reasoning models against both reasoning and non-reasoning ones. The page covers only the reasoning version and notes that a non-reasoning variant may exist. Anthropic has not disclosed the parameter count. Haiku 5.5 takes text and image input, outputs text and has a 1M-token context window.

Haiku 5.5 streams at 243.4 tokens per second but its first token takes 415.30 seconds

On Anthropic's API, Haiku 5.5 at max effort generated 243.4 output tokens per second, ninth of 182 and against a tier median of 110.9. Artificial Analysis measures that speed only while tokens are flowing, after the first chunk has arrived.

The latency figure tells the other half of the story. The time to first token for Haiku 5.5 at max effort came in at 415.30s, against a tier median of 2.17s. Artificial Analysis notes that for reasoning models its time to first answer token includes the thinking the model does before the answer begins.

At max effort Haiku 5.5 costs $0.21 per task and generated 440M tokens

Per token, Haiku 5.5 sits well under its tier medians: $0.10 per million input tokens against $0.25, and $0.50 per million output tokens against $0.90. Cached prompts get a 90% discount. At the 7:2:1 cache-hit, input and output blend that Artificial Analysis uses, the rate works out to $0.08 per million tokens.

Per task, the numbers look different. An average Intelligence Index task cost $0.21 for Haiku 5.5 at max effort, which places it 46th of 182 on cost, against second place on intelligence. Over the full index the model generated 440M output tokens, 4.4 times the 100M median. At max effort, Sonnet 5.5 used 410M tokens on the same index, so the smaller model outtalked its bigger sibling.

Max effort means Haiku 5.5 thinks before every answer, and the thinking is billed as output

According to Anthropic, Haiku 5.5 is the first Haiku with effort controls, so teams can tune cost against intelligence task by task. This page tests the max setting. A reasoning model first writes out a chain of thought, then the answer, and both arrive as output tokens on the bill.

Think of a taxi meter that keeps running while the driver studies the map, not only while the car moves. Fast driving, here 243.4 tokens per second, helps less when the planning runs long, and the planning is charged at the same output rate as the answer.

Anthropic positions Haiku 5.5 for high-volume, cost-sensitive work and for latency-sensitive uses such as chat, voice agents and live support. It also pitches the model as a subagent that a larger model such as Fable or Opus plans around and hands subtasks to, so many agents run in parallel. Anthropic calls it the fastest and most efficient model in the Claude 5.5 family.

Artificial Analysis lists a single price, yet according to Anthropic the $0.10 and $0.50 rates hold only for prompts up to 100K tokens; longer prompts cost $0.50 input and $2.50 output, five times as much. In our view, a 415.30s first-token wait at max effort sits awkwardly with Anthropic's pitch for chat and voice agents, and the page shows no lower effort setting to reveal what Haiku 5.5 gives up to start faster.

Where lower effort settings land

This page covers only the max-effort reasoning configuration, and Artificial Analysis says a non-reasoning variant may also exist. No scores for Haiku 5.5 at lower effort levels appear here, so how the 43 index score and the 415.30s wait change at other settings has not been measured. For now Artificial Analysis lists a single API provider for the model.

Related stories

  1. Sonnet 5.5 hits third place by writing 410M tokens
  2. Haiku 5.5 cuts token prices 90% on prompts under 100k
  3. Sonnet 5.5 edges past Opus 5.5 on Terminal-Bench 4.0
  4. Max effort adds nothing to Fable 5.1's ARC-AGI-2 score
  5. Opus 5 tops Fable 5 on OSWorld 2.0 at a third of the price
  6. Claude Fable 5: access and capabilities

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.