Claude Sonnet 5 handles date ambiguity in HR benchmark

Puter.js tested Claude Sonnet 5 against Sonnet 4.6 using a prompt to calculate employee bonuses based on tenure. The input contained a trap: hire dates were mixed formats. Riley's date (25-04-2019) is DD-MM, while Morgan's (09-30-2022) is MM-DD.
The challenge was Jordan's date (06-07-2021). If read as June 7, tenure exceeds five years, triggering a 7% bonus. If read as July 6, it falls three days short, resulting in a 4% bonus. Sonnet 4.6 completed the task without flagging the conflict, applying both formats inconsistently to reach a result.
Sonnet 5 identified the format mismatch, presented both potential bonus outcomes for Jordan ($100,580 or $97,760), and asked a clarifying question. According to Artificial Analysis, Sonnet 5 has a higher cost per task than Opus 4.8 before accounting for promotional pricing.
Related stories
- Claude Sonnet 5 increases performance at a higher cost per task
- Opus 5.5 aced a test suite at 3.4x GPT-6 Sol's cost
- Opus 5.5 finds new bugs for CodeRabbit and misses old ones
- Fable 5.1 refuses the knife but heats a gas can anyway
- Andon Labs puts GPT-6 Astra ahead of Claude Fable 5.1
- Claude Opus 5 tops Sierra's agent-building test at 23.9%
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
