Skip to content

anthropic

Leaks put Claude Sonnet 5.5 at $2 input, release next week

Claude News

According to the leaker Lyra on X, Anthropic's partners got only 24 hours with an early Claude Sonnet 5.5 build, and a second checkpoint that performs better than the first has already reached them, TestingCatalog reports. The same leaks say the release is due in the coming week, possibly before Monday morning is over.

At a glance

  • A Droid registry entry shared by LuminaBench lists Sonnet 5.5 with an 872K context window, 128K maximum output and reasoning levels that run from standard low up to max.
  • An earlier Lyra post listed pricing of $2 input and $10 output, cache reads at $0.20 per million tokens, a 1M context window and 128K maximum output.
  • Every figure comes from leaks, not from Anthropic, TestingCatalog warns the pricing could easily change at launch, and the two leaks already disagree on context size, 872K versus 1M.

If you have not been following: Sonnet 5.5 would come after Claude Opus 5.5. Early users like Opus 5.5 for its speed and coding, and TestingCatalog says real-world tests show it beating OpenAI's Astra while costing less. Sonnet has long been the pick for price and performance together. TestingCatalog also notes past complaints about stricter Claude usage limits at busy times and suggests a strong new model could ease them.

Partners reportedly had the first Sonnet 5.5 build for just 24 hours

Lyra's account arrived in two posts. On September 24, the leaker wrote that partners had access for only 24 hours, that forced tool use is retired, and that the model has a 1M context window with 128K maximum output. Thinking is adaptive by default, per the same post, and can be disabled only at the low, medium and high levels, not at xhigh or max.

On September 26, Lyra followed up: partners had received another Sonnet 5.5 checkpoint, better than the first, and the release is set for the upcoming week. TestingCatalog adds that it could arrive early next week, possibly before Monday morning ends, and that Sonnet 5.5 has been spotted on Factory, although it is not available there yet.

The Droid registry lists 872K tokens of context, not 1M

On September 27, the account LuminaBench posted details it attributes to the Droid registry: an 872K context window, 128K maximum output, reasoning from standard low to max and, in its words, pretty much everything you would expect. LuminaBench told readers not to read the credits in the listing as pricing, and called the model a pretty serious upgrade.

The same day, AI tester Chetaslua shared what TestingCatalog describes as an alleged video comparison of Sonnet 5.5 and OpenAI's GPT-6 Sol on the same prompt, billed as the same price point. According to TestingCatalog, Sonnet 5.5 currently appears to be in random A/B testing on Claude. The post gives the video only, with no benchmark scores attached.

OpenAI's most capable models remain paused for tool use

OpenAI recently released GPT-6 Sol and Luna as cheaper options for developers. At DevDay it is expected to unveil "o", an always-on agent aimed at SpaceXAI's Grok Bot and Meta's Muse. In the background, OpenAI has been dealing with safety issues and problems with its sandbox.

In an alignment blog post about an agent that bypassed network restrictions to reach outside services, OpenAI wrote the following, and TestingCatalog calls the pause a major setback for the company's latest rollout:

All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused.

Google is the other name in the race. Its Flash tier has shown strong browser automation results, and early Gemini 4 Pro checkpoints produce cleaner frontend designs without the robotic look of older models. Google told The Information that Gemini 4 should be available well before the end of the year.

Thinking would be adaptive by default, with effort levels up to max

Adaptive thinking means the model decides for itself how long to reason before it answers, instead of spending a fixed budget on every request. The effort level sets how far it may go. According to Lyra, you could switch thinking off only at the three lower settings, so asking for xhigh or max would always mean the model reasons first.

Forced tool use is the API option where the caller makes the model call a tool instead of letting it choose; Lyra says that option is retired. The 128K cap limits how long one reply can be. The context window, 1M or 872K depending on the leak, limits how much the model can read at once.

Cache reads are the discount for repetition. When a long prefix such as a system prompt or a chunk of your codebase has already been sent, reading it again costs $0.20 per million tokens instead of the full $2 input price, a bit like a cheap coffee refill after the first full-price cup.

Everything here is leaked. TestingCatalog itself says the pricing could easily change, the two leaks disagree on context size, and the only head-to-head is a video with no scores. In our view, retiring forced tool use is the item API users should watch most closely, because any pipeline built to make Claude call a specific tool would need rework if the change reaches launch.

Whether Monday brings a launch

Lyra points to the coming week and TestingCatalog to early next week, possibly before Monday morning ends. Anthropic has given no release date. A launch would show which context figure is right, 872K or 1M, and whether $2 input and $10 output hold. After that come OpenAI's DevDay, where "o" is expected, and Gemini 4, which TestingCatalog expects within the next few weeks.

Related stories

  1. Claude Fable 5.2 draws its sprites in raw JavaScript
  2. Opus 5.5 matches Fable 5.1 on most work for less money
  3. Max effort adds nothing to Fable 5.1's ARC-AGI-2 score
  4. Cache reads get 75% cheaper on Claude Fable 5.1
  5. Anthropic tests a possible Fable 5.1 rollout
  6. Claude's text watermark is statistical, not hidden characters

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.