ponytail's token savings: 10.3% off the bill, not the promised 20%

JetBrains ran 80 paired tasks against the ponytail skill for Claude Code, the third token-saving add-on in its testing series. The pitch: 54% less code, 22% fewer tokens, 20% lower cost, 27% less time. The median measured result: 15% less code, 10.3% lower cost (p=0.004), and 11% less time.
Setup was Claude Code 2.1.201 with claude-sonnet-5 at medium reasoning effort, using the SkillsBench benchmark. The savings show up where the agent overcomplicates things: code dropped 31% on large builds, while the median didn't budge on tasks that were already short.
One big caveat on installation. Dropping SKILL.md into the skills folder wasn't enough: the model never invoked it once across 10 sessions. Every number here came from a branch where the ruleset is force-injected by a SessionStart hook.
On quality, 65 tasks scored the same, 9 got worse, and 6 got better. The full evaluation run cost $246.09.
Related stories
- rtk promised 60-90% fewer tokens. JetBrains measured 7.6% more spend
- ponytail cuts up to 94% of the code in Claude Code sessions
- By the end of a long session, strict solutions drop to 0.5%
- Claude went from bug hunt to abuse report on New Year's Eve
- Opus 5.5 aced a test suite at 3.4x GPT-6 Sol's cost
- Claude hands paid users a spare limit reset until Oct 22
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
