Claude reviewing Codex: 71.6% to 89.7%. Codex reviewing Claude: 91.4% down to 82.8%.

Researchers ran Claude and Codex over 116 recent hard and medium tasks from the lcb set across six setups: solo runs for both models, both cross-review orders, and both self-reviews. The reviewer got the problem statement and the draft solution but couldn't execute tests, which puts it closer to an ordinary code review.
Claude's review lifted the share of passing Codex drafts from 71.6% to 89.7% (p = 0.001). Codex reviewing itself landed at 84.5% (p = 0.022).
Going the other direction backfired. Codex reviewing Claude's drafts pushed the pass rate down from 91.4% to 82.8% (p = 0.046), and Claude's self-review left its 91.4% baseline untouched.
The preprint, arXiv:2607.21656, went up July 22, 2026, and was accepted at the Agentic SE workshop at KDD'26.
Related stories
- pairmark races Claude Code against Codex in your repo
- Claude Code's /code-review now has effort levels
- Anthropic's models miss the frontier in a security PR-review test
- Claude's PRs merge at 84%, one point under humans
- Claude Code matches Codex on SWE-Bench Pro at twice the cost
- Claude Opus 5 tops Sierra's agent-building test at 23.9%
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
