Skip to content

open-models

Mistral's 1T-parameter Le Chonk keeps only 49B active

Promtime

Mistral's nickname for Mistral Large 4 is "Le Chonk", and the model is large enough to earn it: 1 trillion parameters, of which only 49 billion work on any given request. According to Mistral's announcement, shared on Threads, the API is rolling out today and the open weights are planned for the end of October.

At a glance

  • Mistral Large 4 is a natively multimodal mixture-of-experts model aimed at software engineering, cybersecurity, finance, law, science and manufacturing, and it is available now in public preview through Mistral Studio.
  • The preview costs $1.36 per million input tokens and $4.18 per million output tokens, and Mistral reports 61.7% on DeepSWE v1.1, 93% on Cybench and 42% on Dense 200 visual grounding.
  • The weights stay locked until red-teaming with cybersecurity leaders, vetted partners and state authorities is done, and every benchmark figure so far comes through Mistral's own announcement.

If you have not followed Mistral's naming, the Large line is the company's flagship series. Testing Catalog describes Large 4 as Mistral's largest and most capable model yet. Mistral's own post leans into the size, introducing the model as "Le Chaton Fat" before shortening it to "Le Chonk".

Mistral reports 61.7% on DeepSWE v1.1 and 93% on Cybench

Mistral calls Large 4 the best open-weights model from the US or Europe on aggregated benchmarks, and claims state-of-the-art results on what it calls critical workloads: cyber defense, manufacturing and finance. On software tasks it reports 61.7% on DeepSWE v1.1 and 59.9% on AutomationBench. On security it reports 82% on one Artificial Analysis Cyber Index test covering vulnerability reproduction and patching, and 93% on Cybench.

The vision comparison is much tighter. On Dense 200 visual grounding Large 4 scored 42%, against 41% for GPT-6 Astra. Mistral also says third-party evaluations put it above GPT-6 Astra on legal and financial tasks, that Harvey's Legal Agent benchmark shows it ahead of all open-source models, and that it resisted 93.3% of attacks on Lakera's B3 benchmark.

The preview costs $1.36 per million input tokens and runs in Mistral's own datacenters

The API is open now through Mistral Studio at $1.36 per million input tokens and $4.18 per million output tokens. Large 4 will be offered in several regions, including a European deployment that Mistral operates end to end under European law. Private-cloud and on-premises options are planned for security teams.

Mistral trained the model from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters, the same ones that serve the preview. The training data covered more than 160 languages, including every official language of the European Union. On the input side, Mistral lists spreadsheets, documents, charts, technical drawings, PDFs and large geospatial images.

The open weights wait for red-teaming with partners and state authorities

The weights are due by the end of October, and until then the public preview API is the only way to use Large 4. Before the release, Mistral is red-teaming the model with cybersecurity leaders, vetted partners and state authorities, who work with the same model running with reduced moderation and expanded cyber capabilities.

The release is the first milestone backed by Mistral's €3 billion Series D. Mistral says Large 4 will be the base for a new family of specialized models, pairing open weights with self-deployment and greater customer control. According to Mistral, the base model already combines instruction following, reasoning and agentic capabilities.

Only 49 billion of the 1 trillion parameters work on each token

Large 4 is a mixture-of-experts model. Instead of one dense network doing all the work on every token, the parameters are split into many specialist sub-networks, and a small router picks a handful of them for each piece of input. Total size sets how much the model can hold; the active slice, here 49 billion, sets how much compute each token costs.

Think of a large hospital: it employs hundreds of specialists, but any one patient sees only a few of them, so the hospital can grow without every visit getting slower. Natively multimodal means images and documents go into the same model from the start, with no separate vision add-on bolted on later.

Mistral also described its reinforcement-learning pipeline, the same training, customization and RL environment it offers to enterprises through Mistral Forge. At 3,000 GPUs it generates roughly 33 billion tokens per day, of which about 16 billion are trainable completion tokens after filtering and masking.

Every benchmark here comes through Mistral's announcement, including the third-party results it cites. The announcement names neither the license the weights will carry nor the hardware needed to run them. In our view, the Dense 200 comparison deserves the least weight: 42% against 41% for GPT-6 Astra is a one-point gap on a single test.

When the October weights land

Mistral has named a month, not a day: the weights are due by the end of October, once red-teaming wraps up. It has given no date for the private-cloud and on-premises options, or for the first specialized models built on Large 4. If you plan to self-host a 1-trillion-parameter checkpoint, keep some disk space free.

Related stories

  1. Mistral CEO says new model tops Chinese ones in some areas
  2. Aleph Alpha puts a 1M-context model on your own servers
  3. Ideogram 4.5 prices images from 0.8 to 22 cents at 2K
  4. Xiaomi's new MiMo models carry an unverified top-6 claim
  5. Cohere claims a WMT26 lead with a non-reasoning translator
  6. 552B MoE opens DeepSeek's new architecture family

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.