GLM-5.2 (max) matches Claude Opus 4.8 on Harvey LAB-AA

In an independent measurement by Artificial Analysis, GLM-5.2 (max) reached Claude Opus 4.8's level on all-pass rate in the Harvey LAB-AA benchmark. Harvey's scoring is strict: a task counts only if every rubric criterion passes, with no partial credit.
The benchmark runs on legal work. The model reviews contracts from an M&A deal data room, hunts for change-of-control and assignment provisions, and produces a report for the deal team as coc-analysis-report.docx.
Beyond all-pass rate, Artificial Analysis tracks cost per task broken out by input, cache hit, cache write, reasoning, and answer tokens, plus output tokens per task, decode time, and the average number of agent turns per task.
Related stories
- GLM-5.2 vs. Claude Opus 4.8
- Anthropic's models miss the frontier in a security PR-review test
- Semgrep: GLM-5.2 outperforms Claude Code in IDOR detection
- Qwen 3.8 and Claude Opus 5 show why raw benchmark scores don't predict the bill
- Kimi K3 vs Fable 5: same code, a third of the price, four times slower
- Opus 5.5 aced a test suite at 3.4x GPT-6 Sol's cost
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
