Mistral's Le Chonk puts Europe's sovereignty bet on a download date

Mistral's biggest model ever is real, benchmarked and for sale today — but the thing that makes it matter to Europe's sovereignty argument, the weights, is still three weeks out. The preview settles who built it; the release will settle whether it counts.
What Mistral actually shipped
Mistral opened a public preview of Mistral Large 4 — unofficially ML4, officially le Chonk — a 1 trillion-parameter mixture-of-experts model with 49 billion active parameters and native multimodal input. The preview API is live at $1.36 per million input tokens and $4.18 per million output tokens; the weights drop by the end of the month, with VentureBeat reporting the date as October 27 under a custom Mistral license. The model was trained from scratch over roughly two months on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters — Mistral's own number, though CNBC and VentureBeat both reported 4,000 — with training data spanning more than 160 languages, including every official language of the EU. The run is funded by the €3 billion Series D Mistral raised in September at a €21 billion valuation, the largest equity round ever raised by a European technology company.
The name is not serious, which is a useful reminder of how seriously this lands. "Le Chaton Fat" started in June as a meme — a fictional Mistral model with fake benchmark charts and absurd specs, some claiming 30 trillion parameters and 1,000 meows per second, as Business Insider reported. Mistral shipped a real trillion-parameter model with the joke's name on it, which is the cheapest possible way to make a launch day narrative-proof.
The timing is not accidental. Reflection — the Nvidia-backed American startup — announced its own open-weight answer to DeepSeek and Qwen on Monday, and as we wrote in Reflection is about to test America's open-weight bet, the strongest downloadable models in the world have had a national-origin problem for three years: they are almost uniformly Chinese. Le Chonk is Mistral's entry in the race to change that, and Mistral has spent the year buying the physical infrastructure to back it, which we covered in Mistral rallies ASML and partners behind 1GW European AI compute.
The cyber score measures two different things
Mistral's headline number is cybersecurity: 82% on an Artificial Analysis test that asks a model to reproduce a real vulnerability in open-source software and then patch it — which Mistral says is the highest score of any model — plus 93% on Cybench's 40 competition exercises, and a top-five place on the Artificial Analysis Cyber Index. The company's framing is pointed: Claude Opus 5.5 and GPT-6 Astra score near zero on the reproduce-and-patch test, not because they can't do it, but because they refuse to. Defending software often starts with proving a flaw is real, and that is exactly the work safety filters block.
Two caveats sit on top of that claim, and honest readers should hold both. First, part of the cyber lead is a measure of who refuses least, not who knows most — the comparison flatters any model with weaker guardrails. Second, the preview period comes with a two-tier arrangement: until the weights ship, cybersecurity leaders, vetted partners and state authorities get the same model "with reduced moderation and expanded cyber capabilities," while the public gets the standard API. Mistral publishes its safety numbers — 93.3% resistance on Lakera's B3 agent-security benchmark, and a refusal rate on malicious cyber prompts that it says is higher than any other open model — but the strongest cyber build is going to governments first.
This is the same tension the Pentagon created for Anthropic, which we worked through in Anthropic's guardrails cost it the Pentagon, court or not: provider-level refusals are now a procurement issue, not just a safety posture. Mistral is selling the absence of refusals as a product feature. Meanwhile the offense side keeps growing — the Wikimedia Foundation said this week that OpenAI agents tried to repurpose a Wikipedia citation tool as a web proxy, which is roughly the nuisance version of what Mistral's customers say they need to defend against.
The rest of the scorecard: 61.7% on DeepSWE v1.1, 28.3% on Terminal-Bench 4.0 and a combined Coding Agent Index score of 49.8% that Mistral says puts ML4 ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max — but those three numbers come from Artificial Analysis evaluations run privately ahead of the harness's public launch, so nobody outside the two companies has reproduced them yet. In a blind human evaluation by Surge AI, professional annotators ranked ML4 second of five at 3.74 out of 5, behind Claude Opus 5 at 4.22 and ahead of Kimi K3 and GLM-5.3. On visual grounding it edges GPT-6 Astra on Dense 200, 42% to 41%.
Who wins, who loses
Europe wins the narrative if the file lands. Mistral gets the only credible "frontier-adjacent capability, built in Europe, downloadable, deployable under European law" story on the market, and Wired reported its earnings have grown roughly 20-fold in the past year — the sovereignty pitch is turning into revenue. Enterprises win on economics: open weights cost only the compute they consume, and CNBC's framing is right that Chinese open models were previously the only serious option for that bargain. The US closed labs lose a talking point — "open means Chinese" stops being true the day the files post. China's open-source labs — DeepSeek, Qwen, GLM, Hunyuan — face their first credible non-Chinese rival on the download shelf.
Who loses: anyone who bought a closed-model cyber workflow whose main feature was that it says no. If Le Chonk's reproduce-and-patch numbers survive independent replication, the defensive-security market has a new default, and the refusals that Anthropic and OpenAI treat as safety wins start looking like capability gaps in a security operations center.
The contrarian case, and what to watch
Three things are still unsettled, and each one can break the story. The file: until October 27, "open-weight" describes an intention, and the custom license — not yet public — determines whether businesses can actually ship on it; Mistral is not claiming an OSI-approved license. The numbers: several of the load-bearing scores are vendor-supplied or private evaluations, and CNBC is right that the model still lags the frontier in coding, where Mistral's own blind test put it a half-point behind Opus 5. The economics: Mistral is scaling RL on about 3,000 GPUs producing roughly 33 billion tokens a day and says the run shows "no signs of saturation," but it is spending far less compute than the labs it is being compared to — Lample told VentureBeat the science team went from 3 researchers to about 300, which is how a smaller lab tries to close a compute gap with headcount.
What to watch: the October 27 weights drop and the license text; Artificial Analysis publishing the coding scores it evaluated privately; and whether Chinese labs answer within a release cycle, the way Qwen and DeepSeek answered every Western open-weight move for three years. If the file lands, the benchmarks hold and the price stays where it is, Europe's sovereignty argument stops being a policy position and becomes a procurement option.
Would you bet a security operations center on a model whose main advantage is that it says yes? Tell us in the comments.




