anthropic
Max effort adds nothing to Fable 5.1's ARC-AGI-2 score
Claude News
anthropicARC Prize has published verified ARC-AGI results for Anthropic's Claude Fable 5.1 across five reasoning levels, with max effort scoring 97.5% on ARC-AGI-1 Semi-Private at $1.40 per task and 90.0% on ARC-AGI-2 Semi-Private at $4.49 per task. The results, dated September 1, appear on the ARC Prize results page.
At a glance
- The XHigh setting matches max effort exactly on ARC-AGI-2 Semi-Private, both at 90.0%, while on ARC-AGI-1 the two verified scores separate by a single point, 97.5% at max effort against 96.5% at XHigh.
- Below the top pair the verified scores fall conventionally: High takes 96.0% and 88.8%, Medium 94.5% and 86.3%, and Low drops to 90.0% on ARC-AGI-1 and 78.3% on ARC-AGI-2.
- ARC Prize also published per-task pass and fail grids for all five levels, covering 120 tasks on the ARC-AGI-2 public evaluation set and 400 tasks on the ARC-AGI-1 public set.
The flat top of the ladder is the notable part of this run. Two of the five settings return the same ARC-AGI-2 number, which reads as a limit of the model's reasoning depth on that benchmark rather than a scoring quirk, and it pushes the separation onto ARC-AGI-1, where the two differ by a single point. For teams picking a reasoning level, the verified table is closer to a cost decision than a headline.
The published grids record a pass or fail for each reasoning level on every task. On the ARC-AGI-2 public evaluation set of 120 tasks, 88e364bc fails at all five levels, and the ordering breaks elsewhere: 800d221b fails at max effort while XHigh and High solve it, and 88bcf3b4 is solved only at High.
The 400-task ARC-AGI-1 public set shows the same non-monotonic behaviour in places. Task 0d87d2a6 fails at every level, 8fbca751 and 50f325b5 are each solved at a single setting, High and Medium respectively, and f3b10344 fails at max effort and XHigh while High, Medium and Low all pass.
The ARC-AGI-2 leaderboard lists Claude Fable 5.1 alongside the earlier Claude Fable 5 and Claude Opus 5, together with entries from other labs: Gemini 3.7 Flash, GPT-5.6 Sol, Grok 4.6, DeepSeek V4 Pro 0813, Kimi K3 and Inkling. The Fable 5.1 entry carries no version suffix, while Claude Fable 5 is listed with v1 and v2.
The empty ARC-AGI-3 column The verified table leaves ARC-AGI-3 blank for all five reasoning levels, so the third benchmark generation carries no Claude Fable 5.1 result in this run and no date is given for one. Per-task cost is reported for max effort only: the four lower settings appear with verified ARC-AGI-1 and ARC-AGI-2 scores and no price attached.
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
We only use your name and avatar from Google. We never store your email address.
