Skip to content

benchmarks

AI agents fail most medical tests in the US

Claude News

The CHI-Bench benchmark shows top AI agents struggle with 72% of real-world medical processes. Researchers tested 30 models across 75 complex scenarios, including treatment planning and care management.

The best performer, Claude Code on Opus 4.6, succeeded only 28% of the time. GPT-5.5 from OpenAI came in second at 21%. Reliability drops further on repeat runs; no model passed 20% on three consecutive attempts at a single test.

The testing mirrors real clinic workflows, with dozens of steps. The results suggest AI agents' readiness for autonomous medical work is overstated.

Related stories

  1. Sonnet 5 aligned an early Opus 4.8 checkpoint
  2. Claude Fable 5.1 cracks Urquhart's Cyphral Distich
  3. The last of Wiedijk's 100 theorems falls to Anthropic
  4. Claude's PRs merge at 84%, one point under humans
  5. Anthropic paper: automated researchers fix 10/10 benchmarks
  6. A harness pushes Claude Opus 5 past 30% on ARC-AGI-3

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.