Sonnet 5.5 cut CodeRabbit's review bill by about 60%

Anthropic advertised Sonnet 5.5 as up to 30% cheaper per task than Sonnet 5. In CodeRabbit's code-review evaluation, though, the Claude model calls came out about 60% cheaper per review at identical list prices, and the new model also caught more of the hardest known bugs, 6 of 13 against Sonnet 5's 4, in about half the wall-clock time.
At a glance
- CodeRabbit tested Sonnet 5.5 against Sonnet 5 on 13 hard known-bug pull requests and 44 open-source pull requests, replaying identical recorded inputs so only the model and its thinking setting changed.
- On both benchmarks, Sonnet 5 read more than twice as much input per core review call as Sonnet 5.5, wrote about four times as much output and thought roughly six times as many words.
- Thirteen cases is a small set, Sonnet 5.5 missed two bugs Sonnet 5 caught, Opus 5.5 still caught 8 to 10 of the same 13, and judge scores for the larger run are pending.
If you have not been following the Sonnet line in code review, CodeRabbit has benchmarked every release since Sonnet 4. Sonnet 4.6 was its highest-coverage Sonnet, catching about 63% of known issues, but at 29% precision it commented on everything. Sonnet 5 swapped that for cleaner comments, while coverage fell to about 50% and nitpicks multiplied. For Sonnet 5.5, the open question was whether it could win back coverage and keep the precision.
Sonnet 5.5 caught 6 of 13 hard bugs where Sonnet 5 caught 4
CodeRabbit's Signal set is 13 real pull requests from Elasticsearch, Puma, vLLM, Cilium, axios and Next.js, each with one verified issue a reviewer should catch. Nine are rated difficulty 3, three difficulty 4 and one difficulty 5. An independent judge scored every comment with three votes, and a case counts as caught only when a regular actionable comment wins a majority.
With adaptive thinking on, Sonnet 5.5 caught 6 of 13 (46.2%) at 41.2% actionable precision, with 17 reported comments and two nitpicks. Sonnet 5 caught 4 of 13 (30.8%) at 40.0%, with 15 comments and three nitpicks, and labeled two comments critical, one of them wrongly; Sonnet 5.5 labeled none critical. Four of Sonnet 5.5's catches were bugs Sonnet 5 waved through, while it missed two that Sonnet 5 caught.
Across 44 pull requests, a Sonnet 5.5 review took 6:33 on average against 13:31
The second benchmark, OSS August, holds 85 known issues across 44 open-source pull requests. Sonnet 5.5 with thinking on averaged 6:33 per review against 13:31 for Sonnet 5, with medians of 5:44 and 13:49, and needed 4 h 49 min for all 44 reviews to Sonnet 5's 9 h 55 min. Sonnet 5 also had four generations run longer than ten minutes; Sonnet 5.5 had none.
On OSS August, Sonnet 5.5 posted 111 comments to Sonnet 5's 146, about 24% fewer, with 4 labeled critical against 14 and 9 nitpicks against 30. Judge scoring for this set is pending, so CodeRabbit reads those counts as workload rather than quality. On the 13 Signal reviews, mean time was 5:27 against 9:55, and over both sets Sonnet 5 needed just over twelve hours for 57 reviews where Sonnet 5.5 needed six.
On 44 pull requests, Sonnet 5.5's Claude calls cost $0.46 per review against $1.16
Sonnet 5.5 keeps Sonnet 5's price list: $2 per million input tokens, $10 per million output, $0.20 for cache reads and $2.50 for cache writes, half of Opus 5.5's $4 and $20. So any cost gap between the two Sonnets is purely a token gap. CodeRabbit priced only the Claude model calls, leaving out the shared smaller models that handle summaries and verification.
On the 13 Signal reviews, Sonnet 5.5 with thinking on cost $6.16 in total, or $0.47 per review, against $15.06 and $1.16 for Sonnet 5. On OSS August the totals were $20.32 and $50.95 for 44 reviews. On both benchmarks Sonnet 5.5 came to about 40% of Sonnet 5's cost, which is where the roughly 60% saving comes from.
Opus 5.5 still caught 8 to 10 of the same 13 hard cases
On the identical Signal cases, Opus 5.5 caught 8 of 13 at 66.7% precision in its Standard setting and 10 of 13 at 52.0% at Max, according to CodeRabbit's September evaluation. Anthropic's launch table shows a smaller gap: Sonnet 5.5 lands within about three points of Opus 5.5 on most rows and beats it on Terminal-Bench 4.0, 70.6% against 66.4%, up from Sonnet 5's 10.3%.
CodeRabbit also ran one side-by-side build in Claude Code, a brick-model designer called Brick Studio. Sonnet 5.5 finished in 29 minutes 27 seconds against 44 minutes 50 seconds for Opus 5.5, with near-identical results and slightly higher fidelity from Opus. CodeRabbit calls it one run, not a benchmark.
Sonnet 5 wrote about four times as much per review call as Sonnet 5.5
The saving comes from how much each model reads, writes and thinks. On Signal's core review calls, Sonnet 5.5 averaged 110.7k input tokens, 5.8k output and 464 thinking words; Sonnet 5 averaged 247.5k, 21.6k and 2,771. On OSS August the pattern held: 87.3k, 5.7k and 523 against 191.5k, 23.8k and 3,143.
Picture two colleagues reviewing the same diff. One reads each file once and leaves a short note; the other rereads the files and drafts a long memo first. CodeRabbit says Sonnet 5 behaved like the second, reading the same files repeatedly and writing long deliberations, and across the whole pipeline it used 27% more total tokens than Sonnet 5.5 on Signal and 49% more on OSS August.
Thinking is a separate dial. Sonnet 5.5 allows thinking off at low, medium and high effort, via what Anthropic's migration guide calls between_tools. On Signal, thinking on caught 6 of 13 at 41.2% precision with 17 comments, while thinking off caught 5 at 38.5% with 13 and wrote about 2,900 output tokens per core call instead of 5,800. Thinking on added about 15% to the bill, and CodeRabbit recommends it as the default.
Thirteen cases is a small set, and CodeRabbit says so itself: one of Sonnet 5.5's seven passing comments won on a two-to-one vote, and without it precision drops to 35.3%. Four of the 13 cases defeated every Sonnet configuration, and the two Sonnets disagreed on six. In our view, the 60% saving is the sturdiest number here, because it rests on token counts from 57 reviews, where the catch rates rest on a handful of judge votes.
Waiting on the OSS August verdict. Judge scoring for the 44 OSS August pull requests is still pending, and no date has been given for it. That run will show whether Sonnet 5.5's lower comment volume costs it any coverage at scale. Meanwhile CodeRabbit is moving simple and moderate reviews to Sonnet 5.5 now, with more to follow in the coming weeks.
Related stories
- Opus 5.5 finds new bugs for CodeRabbit and misses old ones
- Sonnet 5.5 edges past Opus 5.5 on Terminal-Bench 4.0
- Sonnet 5.5 hits third place by writing 410M tokens
- Claude helped make claude.ai 3x faster in two weeks
- Opus 5.5 costs less and answers old agent code with 400s
- Opus 5.5 aced a test suite at 3.4x GPT-6 Sol's cost
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
