Skip to content

anthropic

CodeRabbit's reviews on Sonnet 5.5 cost about 60% less

Promtime

In CodeRabbit's review pipeline, Claude Sonnet 5.5 writes about a quarter as much per review call as Sonnet 5 and still finds more bugs, according to CodeRabbit's evaluation. The Claude calls cost $0.47 per review instead of $1.16, and on 13 hard known-bug cases Sonnet 5.5 caught 6 where Sonnet 5 caught 4, at nearly the same precision.

At a glance

  • CodeRabbit ran Sonnet 5.5 through the same pipeline it used for Opus 5.5 earlier this month, and is already moving simple and moderate reviews to the new model, with more to follow.
  • Both models share a price list of $2 input and $10 output per million tokens, so the roughly 60% saving on Claude calls per review comes entirely from Sonnet 5.5 using fewer tokens.
  • Thirteen cases is a small set, the larger 44-PR run has no judge scores yet, and on the same hard cases Opus 5.5 caught 8 (Standard) to 10 (Max) bugs.

CodeRabbit has benchmarked every Sonnet since Sonnet 4, and each release fixed one weakness while exposing another. Sonnet 4.5 added reasoning depth along with a hedging habit. Sonnet 4.6 caught about 63% of known issues but, at 29% precision, commented on everything. Sonnet 5 went the other way: its comments got cleaner, but coverage fell to about 50% and nitpicks multiplied.

Sonnet 5.5 caught 6 of 13 hard bugs, Sonnet 5 caught 4

The Signal set is 13 real pull requests from Elasticsearch, Puma, vLLM, Cilium, axios and Next.js, each with one verified issue. Nine are difficulty 3, three are difficulty 4 and one is difficulty 5. With thinking on, Sonnet 5.5 caught 6 through actionable comments at 41.2% precision and 17 reported comments. Sonnet 5 caught 4 at 40.0% precision with 15 comments.

The extra catches did not come from louder commenting. Sonnet 5.5 posted 2 nitpicks to Sonnet 5's 3 and labeled nothing critical, while Sonnet 5 marked two comments critical and got one of them wrong. The misses also differ. Four of Sonnet 5.5's catches were bugs Sonnet 5 waved through, and Sonnet 5 caught two that Sonnet 5.5 missed.

On 44 pull requests, a review took 6:33 instead of 13:31

OSS August is the broader set: 85 known issues across 44 open-source pull requests. There, Sonnet 5.5 averaged 6:33 per review against 13:31 for Sonnet 5, and 5:44 against 13:49 at the median. Sonnet 5 had four generations that ran longer than ten minutes, and Sonnet 5.5 had none. Across both benchmarks, Sonnet 5 needed just over twelve hours for 57 reviews and Sonnet 5.5 needed six.

On the same set, Sonnet 5.5 posted 111 comments to Sonnet 5's 146, 24% fewer, with 4 critical labels against 14 and 9 nitpicks against 30. Judge scoring for this set is still pending, so CodeRabbit asks readers to treat these counts as workload rather than quality.

Per review call, Sonnet 5 read more than twice as much and wrote four times as much

Sonnet 5.5 lists at $2 input and $10 output per million tokens, the same as Sonnet 5 and half of Opus 5.5's $4 and $20. At those prices, Sonnet 5.5's Claude calls cost $6.16 for 13 Signal reviews and $20.32 for 44 OSS reviews. For Sonnet 5 the totals were $15.06 and $50.95, or $1.16 per review on both sets.

Anthropic's launch figure is up to 30% less per task. CodeRabbit says a review workload, where Sonnet 5 read the same files over and over and wrote long deliberations, goes well beyond that. On an average Signal review call, Sonnet 5 took in 247.5k tokens, wrote 21.6k and produced 2,771 thinking words. Sonnet 5.5 took in 110.7k, wrote 5.8k and thought 464 words.

Picture two colleagues reviewing the same diff. One rereads every file and drafts a long memo before commenting, and the other gets to the point. Across the whole pipeline, including the smaller summary and verification models, Sonnet 5 used 27% more total tokens on Signal and 49% more on OSS August.

Thinking on bought one more catch for about 15% more money

Opus 5.5 rejects explicit thinking toggles, but Sonnet 5.5 accepts thinking off at low, medium and high effort. Anthropic's migration guide calls this setting between_tools. With thinking off, Sonnet 5.5 caught 5 of 13 at 38.5% precision with 13 comments. Thinking on doubled output per core call, from about 2,900 tokens to 5,800, and added about 15% to the bill.

Reviews with thinking on were not slower. Mean time per review was 31 seconds shorter, which CodeRabbit calls within noise. The two settings disagreed in both directions: thinking on caught a vLLM config-context bug and a streaming tool-call serialization case, and thinking off caught an Elasticsearch terms-enum case in a regular comment. CodeRabbit recommends leaving thinking on.

Anthropic's launch notes also carry two warnings for anyone building a pipeline on this model. Sonnet 5.5 follows instructions literally, so "minimize tool calls" is obeyed to the letter. At low effort, it can report a code change as done without running a check unless the prompt asks for one.

The weak spot is sample size. With 13 cases, single judge votes matter: one of Sonnet 5.5's passing comments got through on a two-to-one vote, and without it precision drops to 35.3%. Four cases beat every Sonnet configuration, and Opus 5.5 caught 8 to 10 at 66.7% and 52.0% precision. In our view, the cost result is the sturdier half of the story, since it held on both sets, while the coverage gain rests on two extra catches.

Waiting on the OSS August judges. Judge scoring for the 44-PR set is still pending. It will show whether Sonnet 5.5's lower comment volume costs it coverage at scale, and whether Sonnet 5's extra comments ever turned into extra catches. No date has been given for those results. In the meantime, CodeRabbit is moving simple and moderate reviews to Sonnet 5.5 and plans to move more over the coming weeks.

Related stories

  1. CodeRabbit: Opus 5.5 trades 9 missed bugs for 11 new ones
  2. Agents building agents top out at 23.9%
  3. A $2 model ranks second on Code Arena
  4. Open source harness behind the ARC-AGI-3 jump
  5. Claude Sonnet 5.5 took #3 with 410M tokens of output
  6. Claude Sonnet 5.5 keeps Sonnet 5's price, cuts cost per task

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.