Mistral Large 4 tops open models at just 15% on legal tasks

A 15% score sounds like a failure. Per VentureBeat, though, it's the best open-weight result on Harvey's Legal Agent Benchmark, and it belongs to Mistral Large 4, nicknamed Le Chonk, a 1-trillion-parameter model that outputs text only.
VentureBeat's report puts the same model at 62% on DeepSWE, ahead of GLM-5.3. It also scores 67% on Finch, a set of financial tasks, which is again the top open-weight result. TNW reports that the weights are planned for October 27. When Mistral's CEO spoke to Reuters on announcement day, the remarks named no model and gave no scores.
The catch is that these numbers come from press coverage, and there's no technical report showing how they were measured. Would you trust them before the weights are out?
Related stories
- Mistral CEO says new model tops Chinese ones in some areas
- Xiaomi's new MiMo models carry an unverified top-6 claim
- Cohere claims a WMT26 lead with a non-reasoning translator
- A 125B MoE model previews Qwen4's architecture
- A 4-bit Mac build fits Qwen3.8-27B in 16.1GB
- Zhipu AI will release GLM-5.3 weights in stages after safety reviews
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
