Letting Claude filter its own tool output took Opus 4.6 from 45.3% to 61.6% on BrowseComp

Anthropic published a breakdown of agent harness design: the loop, the tools, context handling, and the guardrails wrapped around the model. Its argument is that harnesses bake in assumptions about what Claude can't do on its own, and those assumptions go stale with every model release.
The numbers back it up. On BrowseComp, giving Opus 4.6 the ability to filter tool output itself by executing code pushed accuracy from 45.3% to 61.6%, and spinning up sub-agents added another 2.8% on top of the best single-agent runs. Given the same context compaction budget, Sonnet 4.5 stalled at 43%, Opus 4.5 hit 68%, and Opus 4.6 reached 84%.
The Pokemon example is the blunt one: after 14,000 steps, Sonnet 3.5 had piled up 31 note files and was stuck in the second city. At the same step count, Opus 4.6 was keeping 10 files organized in folders and had three badges.
Related stories
- Analysis of the agent control model in Claude Code
- Claude Code sessions can now talk to each other
- Claude Code has four kinds of agent loops: /goal, /loop, /schedule
- A hardcoded instruction in Claude Code 2.1.219 and 2.1.220 blocks Opus 5 from spawning sub-agents
- Claude integrates with Lovable
- Anthropic releases Claude Agent SDK for building AI agents
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
