Skip to content

anthropic

Short prompts to Claude Haiku 5.5 cost 90% less

Promtime

Haiku 4.5 scored 0.0% on Terminal-Bench 4.0, a test of multi-step command-line work, while Claude Haiku 5.5, which Anthropic announced on its Haiku page on October 7, scores 39.2%. It also charges $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100K tokens, a tenth of Haiku 4.5's $1/$5.

At a glance

  • Haiku 5.5 is available on every Claude.ai plan from Free to Enterprise, in Claude Code, on the Claude API as claude-haiku-5-5, and through AWS, Google Cloud and Microsoft Foundry.
  • Above 100K tokens the price rises to $0.50/$2.50, and The New Stack reports that Anthropic puts average savings over Haiku 4.5's flat $1/$5 at around 75%.
  • Haiku 5.5 is the first Haiku with effort controls, defaulting to medium according to The New Stack, though its updated tokenizer uses slightly more tokens per task, which trims the headline savings.

If you haven't been following: Haiku 4.5 arrived on October 15, 2025, matching Sonnet 4 on coding, computer use and agent tasks with 73.3% on SWE-bench Verified. It then stayed the current Haiku while Sonnet 5 (June 30, 2026), Opus 5 (July 24), Opus 5.5 (Sept 22) and Sonnet 5.5 (Sept 28) shipped around it. The New Stack calls Haiku 5.5 the first new version of Anthropic's smallest model in nearly a year.

Above 100K tokens, Haiku 5.5 costs $0.50/$2.50, half the old rate

On the Claude Platform, the price of Haiku 5.5 depends on prompt length: $0.10/$0.50 per million input/output tokens up to 100K tokens, and $0.50/$2.50 above that. The New Stack notes that Haiku 4.5 costs the same however many tokens a request carries, and puts the per-token cuts at 90% and 50% respectively.

According to The New Stack, Anthropic says about 90% of requests to Haiku 4.5 fell into the cheaper under-100K bucket. Anthropic puts the average savings at around 75%. That estimate accounts for the mix of requests and for an updated tokenizer that uses slightly more tokens per task.

The usual discounts still apply on top: up to 90% off with prompt caching and 50% with batch processing. For comparison, Sonnet 5.5 costs $2/$10 per million tokens, Opus 5.5 costs $4/$20 and Fable 5.1 costs $10/$50. The New Stack also notes that decision models like Jev already handle high-volume work at an even lower price.

Terminal-Bench 4.0 went from 0.0% to 39.2% in one Haiku generation

The comparison table published by The New Stack lists four models in this order: Haiku 5.5, Haiku 4.5, OpenAI's GPT-6 Luna and Sonnet 5.5. On Terminal-Bench 4.0, which measures agentic coding, they score 39.2%, 0.0%, 16.4% and 70.6%. On the offline subset of OSWorld 2.1, a computer-use test, they score 72.4%, 15.7%, 48.9% and 83.9%.

The knowledge-work rows follow the same pattern. GDPval-AA v2.1 gives 1620, 735, 1437 and 1840, and AA-Briefcase v1.1 gives 1578, 614, 1336 and 1824. On Chartography (no tools), a visual reasoning test, the four score 46.4%, 6.4%, 29.1% and 61.6%.

Humanity's Last Exam has no GPT-6 Luna entry. Haiku 5.5 scores 45.9% without tools and 57.4% with them, against 10.2% and 18.7% for Haiku 4.5 and 56.9% and 64.5% for Sonnet 5.5. FrontierCode 1.1 (Main) has no Haiku 4.5 entry: Haiku 5.5 posts 46.4%, GPT-6 Luna 42.4% and Sonnet 5.5 52.1% (Xhigh).

HubSpot measured 92.8% on its CRM suite, its best score yet

HubSpot evaluates models on simulated CRM portals. It says Haiku 5.5 got the best score it has seen on that suite, 92.8% averaged over three runs. One audit task asks models to find stale but ambiguous records. Of all the models HubSpot tested, Haiku 5.5 finished that task fastest and had the highest hit rate and the lowest false positive rate.

Box says that in early testing Haiku 5.5 scored 11 points higher than Haiku 4.5 at about half the latency. Another customer runs an Ask in Document feature at about 8M calls a week. On 400 queries, Haiku 5.5 scored 0.84 against 0.76 for Haiku 4.5, which that customer calls a statistically significant improvement.

A team testing its AI Teammates agent product measured over 30% lower latency for task completions and up to 2.5x faster inference per agent turn, compared with the model it uses today. With Haiku 5.5 as the sidekick, Devin Fusion holds a FrontierCode score of 66.2 while cutting cost and latency. The Devin CLI pairs it with Opus 5.5 as the lead.

Anthropic pitches Haiku 5.5 as a subagent under Fable or Opus

The main pattern Anthropic describes is delegation. A larger model such as Fable or Opus plans the work and hands well-defined subtasks to Haiku, so many agents can run in parallel. One quoted customer gives an example: while a bigger model builds a slide deck, a Vanilla subagent reads a 10-K and pulls the segment revenue line the deck needs.

Anthropic also lists summarization, classification, routing and compaction. It names real-time chat, voice agents and live support, repetitive computer-use jobs like form filling and data entry, and small edits that need to apply across many files. For complex coding and knowledge work, Anthropic calls Opus 5.5 the daily driver and Sonnet 5.5 a good fit for well-scoped tasks.

Effort controls come to Haiku for the first time, with medium as the default

Haiku 5.5 is the first Haiku with effort controls, which set how many tokens the model spends on a task, and The New Stack reports that the default is medium. According to the Claude Platform Docs, Haiku 4.5 used manual extended thinking with a fixed budget_tokens value, while Haiku 5.5 is listed with adaptive thinking.

Effort works like the cycles on a dishwasher: a quick rinse for a coffee cup, the heavy cycle for a roasting pan, and you pay for the water accordingly. Teams set it per task to trade cost against intelligence.

Anthropic gives no SWE-bench Verified score for Haiku 5.5 to set against Haiku 4.5's 73.3%, and never quantifies the tokenizer's extra tokens per task, so the 75% average is hard to check against your own traffic. In our view, the 100K threshold is the choice to watch: the Claude Platform Docs list a 1M-token window, and any prompt past a tenth of it pays five times the base rate.

Where Fable 5.5 stands

The New Stack counts Haiku 5.5 as Anthropic's third 5.5 model in a month and says Fable 5.5, the one still missing, will likely go through a much longer review process. No release date has been given. Until Anthropic says how many extra tokens the new tokenizer adds, your real savings depend on how much of your traffic stays under 100K tokens.

Related stories

  1. Opus 5.5 lands 40% cheaper to run than Opus 5
  2. Cache reads drop 75% with Fable 5.1
  3. Claude Max subscribers get up to $200 a month for the API
  4. Google's Nano Banana 2.1 costs $0.076 per 4K image
  5. GPT-6.1 Sol matches Astra on DeepSWE at a fifth of the cost
  6. Claude Sonnet 5.5 took #3 with 410M tokens of output

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.