Skip to content

anthropic

Claude Sonnet 5 handles date ambiguity in HR benchmark

Claude News

Puter.js tested Claude Sonnet 5 against Sonnet 4.6 using a prompt to calculate employee bonuses based on tenure. The input contained a trap: hire dates were mixed formats. Riley's date (25-04-2019) is DD-MM, while Morgan's (09-30-2022) is MM-DD.

The challenge was Jordan's date (06-07-2021). If read as June 7, tenure exceeds five years, triggering a 7% bonus. If read as July 6, it falls three days short, resulting in a 4% bonus. Sonnet 4.6 completed the task without flagging the conflict, applying both formats inconsistently to reach a result.

Sonnet 5 identified the format mismatch, presented both potential bonus outcomes for Jordan ($100,580 or $97,760), and asked a clarifying question. According to Artificial Analysis, Sonnet 5 has a higher cost per task than Opus 4.8 before accounting for promotional pricing.

Related stories

  1. Claude Sonnet 5 increases performance at a higher cost per task
  2. Opus 5.5 aced a test suite at 3.4x GPT-6 Sol's cost
  3. Opus 5.5 finds new bugs for CodeRabbit and misses old ones
  4. Fable 5.1 refuses the knife but heats a gas can anyway
  5. Andon Labs puts GPT-6 Astra ahead of Claude Fable 5.1
  6. Claude Opus 5 tops Sierra's agent-building test at 23.9%

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.