Models & Research

Mistral’s Le Chonk activates 49B of its 1T parameters on 3,800 GPUs

Mistral launched a public preview of Mistral Large 4 on Tuesday, and its own announcement page carries two numbers that the coverage left out. The model activates 49 billion of its 1 trillion parameters, and Mistral says it trained on 3,800 Nvidia Grace Blackwell GPUs, not the 4,000 that three write-ups reported.

The model is nicknamed le Chonk, which is where the parameter count comes in. Mistral describes it as “a 1 trillion-parameter natively multimodal model with 49 billion active parameters” and as its largest and most capable to date. For now you can only reach it through the preview API on Mistral Studio. The announcement says weights drop at the end of the month, and ZDNet reported that the company gave 27 October as the date.

The GPU count three outlets rounded up

The compute figure is where the sources split. CNBC reported 4,000 Grace Blackwell GPUs over two months, and TechCrunch quoted VP of Science Pierre Stock calling it “only 4,000 Nvidia GPUs,” two to three times less than Mistral’s Chinese competitors. ZDNet printed 4,000 as a direct quote from the company, so the number came from the briefing. The published page says 3,800.

That matters because the compute efficiency claim is the pitch. Even so, a 200-GPU difference doesn’t change the story, and it’s the kind of detail that gets repeated for a year once three outlets agree on it. Separately, the announcement puts Mistral’s reinforcement learning run at 3,000 GPUs, producing roughly 33 billion tokens a day, of which about 16 billion are trainable after filtering.

Detailmistral.ai announcementWhat the write-ups ran
Training GPUs3,800 Grace Blackwell4,000 (CNBC, TechCrunch, ZDNet)
Active parameters49 billion of 1 trillionnot reported
Weights release“by the end of the month”27 October (ZDNet)
Series D valuationnot stated$24B (Wired), €21B (CNBC), $24.39B (TechCrunch)
Preview price per M tokens$1.36 input, $4.18 outputnot reported
Figures as published on 6 October 2026. Percentage of active parameters is our calculation.

The bigger omission, though, is the active parameter count. At 49 billion active out of 1 trillion total, roughly 4.9 percent of the model runs on any given token, by our calculation. Mistral’s product listing calls it a hybrid instruct-and-reasoning mixture of experts, which means that sparsity is by design. It also means the “1 trillion” in every headline is the total parameter count, not the number that runs on each query.

Where the announcement is most specific is cybersecurity, but almost none of it reached print. Mistral says ML4 ranks in the top five globally on the Artificial Analysis Cyber Index, scores 82 percent on a test that reproduces a real vulnerability and then patches it, and solves 93 percent of the 40 exercises in Cybench. It also claims 93.3 percent attack resistance on Lakera’s public B3 benchmark. The reason it tops that patching test is not capability alone.

Several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the same test because they refuse to perform the task.

Mistral, Mistral Large 4 announcement

The coding claim is softer than the cyber one

The catch is coding, and that is where CNBC wrote the model still lags the frontier. Mistral’s own numbers agree. The company ran a blind human evaluation with Surge AI in which ML4 Preview placed second of five models at 3.74 on a 1 to 5 scale. It finished ahead of GLM-5.3 at 3.60, Kimi K3 at 3.59 and GLM-5.2 at 3.40, and behind Claude Opus 5 at 4.22.

On the benchmarks it chose to publish, Mistral reports 61.7 percent on DeepSWE v1.1, 59.4 percent on SWE-Atlas-QnA and 28.3 percent on Terminal-Bench 4. That gives a combined Coding Agent Index score of 49.8 percent, which the company says puts it ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max.

One more figure thins out at the primary, and it’s the money. Mistral describes the round behind this work as a “€3 billion Series D,” the largest equity round ever raised by a European technology company, and attaches no valuation to it at all. Wired put the raise at $3.3 billion and a $24 billion valuation, CNBC at $3.4 billion and €21 billion, and TechCrunch at about $24.39 billion. We reported the round as €3 billion at a €21 billion valuation when Samsung led it in September.

Until the weights land, Le Chonk is a preview endpoint with reduced moderation offered to vetted partners and state authorities, which is a narrower thing than the freely downloadable model that some coverage implied. Z.ai delayed GLM-5.3’s open weights over cyber risk earlier this year, so a three-week red-teaming gap isn’t unusual. The date to watch is 27 October, and the number to check against the write-ups is 3,800.

Get the daily rundown

One email each weekday with the AI news that matters, every claim linked to its primary source.

Free, one email each weekday, unsubscribe in one click. We never sell or share your address.

Rundowns AI Desk

Rundowns AI Desk covers artificial intelligence: model releases, research, funding and policy. Every story is written from primary sources, with each claim linked to the announcement, filing or paper it came from, and checked against those sources before publication.

Leave a Reply

Your email address will not be published. Required fields are marked *