AI agents fail most medical tests in the US

The CHI-Bench benchmark shows top AI agents struggle with 72% of real-world medical processes. Researchers tested 30 models across 75 complex scenarios, including treatment planning and care management.
The best performer, Claude Code on Opus 4.6, succeeded only 28% of the time. GPT-5.5 from OpenAI came in second at 21%. Reliability drops further on repeat runs; no model passed 20% on three consecutive attempts at a single test.
The testing mirrors real clinic workflows, with dozens of steps. The results suggest AI agents' readiness for autonomous medical work is overstated.
Related stories
- Sonnet 5 aligned an early Opus 4.8 checkpoint
- Claude Fable 5.1 cracks Urquhart's Cyphral Distich
- The last of Wiedijk's 100 theorems falls to Anthropic
- Claude's PRs merge at 84%, one point under humans
- Anthropic paper: automated researchers fix 10/10 benchmarks
- A harness pushes Claude Opus 5 past 30% on ARC-AGI-3
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
