Claude Haiku 5.5 takes 415.30 seconds to say its first word

On Anthropic's API, the max-effort version of Claude Haiku 5.5 takes 415.30 seconds, close to seven minutes, to deliver its first answer token, according to Artificial Analysis. The median reasoning model in its price tier starts answering in 2.17 seconds, yet Haiku 5.5 scores 43 on the firm's Intelligence Index, #2 of 182 models in its class.
At a glance
- Anthropic released Claude Haiku 5.5 on October 7, 2026, and Artificial Analysis has profiled its reasoning version at max effort, which takes text and image input and has a 1M-token context window.
- On Anthropic's API, pricing is $0.10 per million input tokens and $0.50 per million output tokens, below class medians of $0.25 and $0.90, and output streams at 243.4 tokens per second.
- The catch is verbosity: running the index, Haiku 5.5 at max effort produced 440M output tokens against a median of 100M, which ranks it #62 of 182 on that measure.
If you haven't been following: according to ClaudeKit, Anthropic's Opus, Sonnet and Haiku tiers stand for most capable, balanced and fastest, and the Claude 5.5 family arrived top down, with Opus 5.5 on September 22, 2026, Sonnet 5.5 on September 28, 2026, and Haiku last. Anthropic has long pitched Haiku as a small model that matches the previous generation's larger ones; it said Haiku 4.5, released October 15, 2025, matched Sonnet 4 on coding, computer use and agent tasks.
Haiku 5.5 scores 43 on an index where its class median is 13
Version 4.3.2 of the Artificial Analysis Intelligence Index combines 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1. The page labels them as agentic knowledge work, SaaS workflows, terminal coding, physics reasoning, document reasoning and long-context reasoning.
Haiku 5.5 at max effort scores 43 and ranks #2 of 182 models in its comparison class, where the median is 13. The class is set by price, not size. Proprietary models are compared with proprietary and open-weights models in the same price band, using a blended 3:1 input-to-output price, and the bands are under $0.15, $0.15–$1 and over $1 per million tokens.
By that formula, Haiku 5.5 blends to $0.20 per million tokens, which places it in the middle band. Reasoning models like this one are compared against both reasoning and non-reasoning models, not against reasoning models alone.
At $0.10 per million input tokens, Haiku 5.5 undercuts its class median
Anthropic's API charges $0.10 per million input tokens and $0.50 per million output tokens, against class medians of $0.25 and $0.90. Cached prompts get a 90% discount. Using Artificial Analysis's 7:2:1 blend of cache hits, fresh input and output, the rate works out to $0.08 per million tokens.
The cost per task is less impressive. Running the index cost $0.21 per task on average, which ranks Haiku 5.5 #46 of 182 on cost, a long way behind its #2 on intelligence.
According to Anthropic, the price also depends on prompt length. The $0.10 and $0.50 rates apply to prompts up to 100K tokens, while longer prompts cost $0.50 for input and $2.50 for output per million. Anthropic also cites savings of up to 90% with prompt caching and 50% with batch processing.
Haiku 5.5 waits 415.30 seconds, then streams 243.4 tokens per second
Once Haiku 5.5 starts answering, it is quick: 243.4 output tokens per second on Anthropic's API, #9 of 182, against a class median of 110.9. Artificial Analysis measures that speed only after the first chunk arrives, so it tells you nothing about the wait before it.
That wait is long. Time to first token is 415.30 seconds on Anthropic's API, while the median for reasoning models in the same price tier is 2.17 seconds. For reasoning models, Artificial Analysis includes thinking time in that figure.
The tokens behind the wait appear in the verbosity figure. While running the index, Haiku 5.5 at max effort generated 440M output tokens, against a median of 100M, which ranks it #62 of 182 on verbosity. Artificial Analysis splits that output into answer tokens and reasoning tokens.
Why does a model this fast take seven minutes to start?
Because it thinks first. A reasoning model writes a long hidden draft, working through the problem step by step, before it produces the answer you see. Picture a fast typist who fills a scratchpad for seven minutes before touching the keyboard: the typing speed is real, but you only see it at the very end.
According to ClaudeKit, Anthropic's thinking modes go back to Claude 3.7 Sonnet in February 2025, its first hybrid reasoning model, and Opus 4.7 later added adaptive thinking. Anthropic's Haiku 5.5 page says this is the first Haiku with effort controls, which let teams trade cost against intelligence on each task. The Artificial Analysis page measures the Max setting.
Anthropic pitches Haiku 5.5 for latency-sensitive work such as chat, voice agents, live support and in-app assistants. It also pitches it as a cheap subagent: a larger model like Fable or Opus plans the job and hands subtasks to many Haiku instances running in parallel.
Artificial Analysis shows only the Max configuration of Haiku 5.5
Several things are missing from the page. Artificial Analysis marks the per-benchmark charts for this model as not publicly available, and Anthropic has not disclosed the model size or parameter count. The page covers only the Max configuration, notes that a non-reasoning variant may also exist, and lists the model with a single API provider.
In our view, a 415.30-second first token at max effort sits awkwardly with Anthropic's pitch of voice agents and live support. At that setting, Haiku 5.5 likely suits batch and subagent jobs better than live conversation.
What the lower effort settings show
The open question is how latency, score and cost change when you turn effort down. Anthropic says its effort controls let teams tune cost against intelligence per task, but this page has no figures for lower settings or for the non-reasoning variant, and no date has been given for when such measurements will appear.
Related stories
- Claude Sonnet 5.5 took #3 with 410M tokens of output
- Short prompts to Claude Haiku 5.5 cost 90% less
- CodeRabbit's reviews on Sonnet 5.5 cost about 60% less
- Claude Sonnet 5.5 keeps Sonnet 5's price, cuts cost per task
- CodeRabbit: Opus 5.5 trades 9 missed bugs for 11 new ones
- First OpenAI win on Vending-Bench, at almost 3x the cash
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
