34% to 44%: the gain came from the conversation, not the memory files

Andrew Jesson fed Claude Code (Opus 5, xhigh effort) a stream of tasks from a simulated business process plus one instruction: get better. Success on held-out tasks climbed from 34% to 44%.
Over the run the agent made 84 writes and edits, every one of them in memory files. It never ran code, never spawned subagents, never used search. The notes formed a graph with correct cross-references. But across 120 held-out tasks, the files themselves were worth +1.7 percentage points, and the confidence interval includes zero.
What actually moved the number was the conversation. Logging completed tasks along with the scores they received is worth +9.9 points, and a summary written after context compaction is worth +5.4 points at a cost of $0.14 versus $6.33 for the full conversation.
A session starting from scratch solves 41 of 120 tasks in 29.8 turns and $0.86.
Related stories
- 34 hours of agent time, about 40 minutes on the actual research question
- Claude Code matches Codex on SWE-Bench Pro at twice the cost
- Sonnet 5 aligned an early Opus 4.8 checkpoint
- Claude Fable 5.1 cracks Urquhart's Cyphral Distich
- The last of Wiedijk's 100 theorems falls to Anthropic
- Claude's PRs merge at 84%, one point under humans
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
