anthropic
Opus 5 tops a new conceptual reasoning index at 73.6
Claude News
anthropicRedwood Research and Anthropic have released the Conceptual Reasoning Index, an aggregate benchmark on which Opus 5 leads with 73.6 points against an estimated ceiling of around 91. The results, dated August 10, 2026, were published by Anthropic alongside three underlying datasets.
At a glance
- The index covers questions with no verifiable answer, where a model must rely on argumentation: reasoning about future AI systems, governance choices and values that empirical feedback loops cannot settle.
- LMCA contributes 60% of the score, ACCoRD and DTBench capabilities 20% each, and the three datasets hold 560 position texts, 567 vetted consistency constraints and 407 decision theory questions.
- Scores have risen roughly linearly since late 2024 with no flattening, and the authors expect LMCA to begin saturating in about a year while DTBench capabilities is already near its ceiling.
Why it matters Benchmarks that reward verifiable answers have shaped what frontier models are good at, and the CRI is an attempt to put a number on the opposite category. For teams that use Claude for research planning or safety analysis, a published gap between 73.6 and an estimated ceiling of 91 reads as a caution about how far argument-heavy output can currently be trusted. The linear trend line likely makes the index more useful as a tracker than as a verdict.
LMCA measures models against expert ratings of 1,461 counterarguments
LMCA, short for Language Model Conceptual Argumentation, contains 560 position texts and 1,461 arguments against them, covering decision theory, philosophy and risks from advanced AI. Nearly all arguments were rated by Redwood researcher Emery Cooper, with some rated independently by others, for 2,140 ratings in total.
Models are scored on how closely their ratings, which follow a detailed rubric, match the human ones. A validation set of roughly 50 arguments was rated independently by four to six people and then discussed for seven to eight hours; inter-rater agreement among humans is high relative to human-model agreement.
The dataset also supports scoring argument generation: one model writes a new argument against a position text, and another model, few-shot prompted with existing rated arguments and the rubric, grades it. Only judging performance currently feeds the index; the authors hope to add generation later.
ACCoRD contributes 567 vetted consistency constraints out of nearly 14,000 generated
ACCoRD, the Assessment of Consistency in Conceptual Reasoning Domains, asks separate model instances for probability estimates or preference orderings and checks whether the answers hold together, for example whether a reported P(A) is at least as large as P(A&B). The dataset holds close to 14,000 model-generated constraints across 18 types, filtered by an automated checker pipeline.
Only the 567 constraints the team checked and approved by hand enter the CRI. The authors treat inconsistency on a set of questions as a signal that a model's reasoning there cannot be trusted by default, and broad inconsistency as evidence of weak conceptual reasoning.
DTBench capabilities is 407 handcrafted multiple-choice questions on decision-theoretic situations involving predictions of a model's own behavior or interactions with near copies. Most were written by Caspar Oesterheld and validated by Cooper. A further 130 questions measuring decision-theoretic attitudes sit outside the index.
Opus 5 scores 73.6 on the index, with DTBench capabilities already close to its ceiling
All models were run at their maximum token limits and effort levels. Opus 5 leads the index at 73.6 points (95% CI ±2.1). Zero corresponds to random guessing and 100 to the highest possible score across all benchmarks, though the authors put realistic ceiling performance near 91.
A model giving maximally good LMCA ratings would score roughly 85 rather than 100, based on expert inter-rater agreement, while a perfect DTBench capabilities run would land at or near 100 and full consistency on ACCoRD yields 100 on that component.
Claude Fable 5 answers 98% of DTBench capabilities questions correctly; its overall CRI score was computed using Opus 5 as a fallback where Fable 5 refused a question. GPT-4's ACCoRD figure rests on incomplete data, since the model refused to fully answer 18% of the benchmark's items. Claude Fable 5, Muse Spark 1.2 and Gemini 3.6 Flash are charted as their companies' leaders on many external benchmarks, though not on the CRI.
What's next The team plans to add benchmarks, retire saturated ones and possibly adjust the weights, and says the site conceptualreasoning.ai will track new models and benchmarks as they appear. Access to LMCA runs through a request form. When ACCoRD will saturate is something the authors describe themselves as very uncertain about.
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
We only use your name and avatar from Google. We never store your email address.
