Context compression cut 39% of tokens and saved almost nothing

One developer ran headroom against SR SWE bench in Harbor: 25 tasks, three branches, plain Claude versus two compression modes, token mode and cache mode. The tasks were picked to favor the tool and everything was billed at Anthropic's API rates. Input tokens dropped from 215M to 132M, a 39% cut. The money came out to $0.37 per task with a negative median, and 13 of the 25 tasks got more expensive.
The cache explains it. On the baseline branch roughly 98% of input tokens were served from cache at $0.30 per million, against $3.00 for fresh ones. Because compression rewrites the history, token mode invalidated the cache 123 times versus 14 for cache mode, pushing 6,062,098 tokens to full price. The lost discount: $16.37, against $15.41 saved.
headroom's own counter overstated the savings by about 1.9x. Solve rates barely moved: 5 of 25 for token mode, 4 each for the baseline and cache mode.
Related stories
- rtk promised 60-90% fewer tokens. JetBrains measured 7.6% more spend
- ponytail cuts up to 94% of the code in Claude Code sessions
- Claude hands paid users a spare limit reset until Oct 22
- Anthropic deleted the post that called a cut a 25% raise
- Claude Code stops billing API users for the auto mode check
- An undocumented Claude Code command outlives the 5-hour cap
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
