Skip to content

claude-code

34% to 44%: the gain came from the conversation, not the memory files

Claude News

Andrew Jesson fed Claude Code (Opus 5, xhigh effort) a stream of tasks from a simulated business process plus one instruction: get better. Success on held-out tasks climbed from 34% to 44%.

Over the run the agent made 84 writes and edits, every one of them in memory files. It never ran code, never spawned subagents, never used search. The notes formed a graph with correct cross-references. But across 120 held-out tasks, the files themselves were worth +1.7 percentage points, and the confidence interval includes zero.

What actually moved the number was the conversation. Logging completed tasks along with the scores they received is worth +9.9 points, and a summary written after context compaction is worth +5.4 points at a cost of $0.14 versus $6.33 for the full conversation.

A session starting from scratch solves 41 of 120 tasks in 29.8 turns and $0.86.

Related stories

  1. 34 hours of agent time, about 40 minutes on the actual research question
  2. Claude Code matches Codex on SWE-Bench Pro at twice the cost
  3. Sonnet 5 aligned an early Opus 4.8 checkpoint
  4. Claude Fable 5.1 cracks Urquhart's Cyphral Distich
  5. The last of Wiedijk's 100 theorems falls to Anthropic
  6. Claude's PRs merge at 84%, one point under humans

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.