benchmarks
One model cleared all five test categories at 100%
Promtime
benchmarksglm-5.3 became the first model on the Ed-o-meter leaderboard to pass all five test categories at 100%, at $0.28 for a full 28-task lap. The leaderboard, published by Reinvently and built by Ed Yau, Applied AI Architect at Kerv, ran 17 models through the same tasks on one harness with verbatim prompts and a single trial each.
At a glance
- The suite splits into 7 coding, 4 data, 9 real-world, 6 security and 2 tool-use tasks, graded by binary automated checkers plus one LLM rubric, with latency measured over a single OpenRouter streaming path.
- Speed is the trade-off for the top finisher: glm-5.3 posts a 16.3-second median time to first token, against 13.2 seconds for gpt-5.5, which cost $1.43 for the same lap.
- Two Anthropic models hit a provider-side classifier: opus-5 had four benign coding-debug tasks blocked before a token was generated, and fable-5 refused five tasks, dropping it to 79% overall.
A single-trial suite of 28 tasks is not a definitive ranking, and the leaderboard says so with wide Wilson intervals. What it does offer is a like-for-like cost and latency picture across vendors, which is the part most internal evaluations skip. The classifier blocks on two Anthropic models look like a measurement hazard rather than a capability gap, and they show how provider-side safety layers can distort a benchmark that scores refusals as failures.
glm-5.3 clears all five categories at 100% for $0.28 a lap
glm-5.3 is the first model on the board to clear coding, data, real-world, security and tool-use at 100%. It backs that with a 9.3 rubric score, third-highest on the board, and $0.28 for the full 28-task lap, roughly a fifth of gpt-5.5's cost. The one price is patience: a 16.3-second median time to first token.
gpt-5.5 is the faster alternative at 13.2 seconds, with the same 100% security score but an 89% real-world corner and $1.43 for the lap. Yau advises checking with compliance before standardising on glm-5.3, and picks sonnet-4-6 for quality without the wait.
kimi-k3 tops the rubric at 9.5 with a 26.4-second time to first token
kimi-k3 holds the highest rubric score on the board at 9.5, judged independently by fable-5, alongside a 96% pass rate. Its weak corner is data development at 75%, and its 26.4-second median time to first token is the second slowest, which the leaderboard rules out for interactive use.
gpt-5.6-luna runs the whole lap for $0.064, or $0.0023 per task, at a 5.3-second median time to first token. Against that it posts 79% overall and 33% on security, which the write-up limits to high-volume, low-risk work where failures are cheap to detect and retry.
haiku-4-5 passes 96% at $0.0044 per task with a 0.9-second median time to first token. deepseek-v4-pro matches that 96% pass rate at $0.0029 per task, but its 40.0-second median time to first token is the slowest on the board, leaving it as a batch-only option.
The gpt-5.6 trio emitted the jailbreak canary in 11 of 12 cells
gpt-5.6-luna, gpt-5.6-terra and gpt-5.6-sol emitted the jailbreak canary in 11 of 12 jailbreak cells, for security pass rates of 33–50%. The three Claude models on the board went 6/6 clean on the security tasks, as did gpt-5.5.
opus-5 posts the best rubric on the default panel at 9.4 and 100% on both real-world and security, then shows 43% on coding. Four benign coding-debug tasks were blocked by a provider-side classifier before a token was generated, on a set of tasks overlapping those blocked on fable-5.
The leaderboard treats that as a measurement hazard rather than a model quirk, and notes opus-5 was penalised twice for flagging an attack it had resisted. Its headline run cost of $1.67 is the true all-trials total, with the blocked trials billed at $0.
Pending an independent re-judge
fable-5's 9.3 rubric is self-judged, and its judge-bias matrix shows it rating itself 9.3 against 8.6–8.7 for models it judges independently. That figure also comes from the earlier 5 July 2026 run, with 11 of 28 trials judged, and is shown pending an independent re-judge that has no date.
The 23 August 2026 change log adds glm-5.3, grok-4.6, deepseek-v4-pro and gemini-3.7-flash. Harness, tasks and checkers are open source at Featherbench under MIT, the full suite costs about $30 to run, and new models can be requested by GitHub issue.
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
We only use your name and avatar from Google. We never store your email address.
