Mistral Large 4: Le Chonk Doesn’t Settle Sovereignty. It Settles the Bill
What was announced
Facts first, no adjectives. Mistral Large 4 is a mixture-of-experts model with 1.05 trillion parameters, 52 billion of them active per token, a vision encoder, and a one-million-token context window. Its predecessor, Large 3, released in December 2025, topped out at 675 billion. The new model was trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in Mistral’s datacenters “in Europe” (the company doesn’t say which country), and the preview runs on that same infrastructure.
It’s available as an API preview at a list price of $1.36 per million input tokens and $4.18 per million output tokens, currently cut in half by a promotion. The weights are due by the end of the month.
This is the first flagship model to come out of the in-house datacenter strategy, the very one Mistral took on 830 million in debt to finance. That alone makes the announcement matter.
What I wrote on September 23, and what has changed
My argument fit in a single sentence: by hosting third-party models, starting with China’s GLM, Mistral was drifting from research lab to platform, and reducing sovereignty to a European access point. Le Chonk seems to prove me wrong. A company that had become nothing more than a storefront wouldn’t train a trillion-parameter model on its own machines.
Except that all it takes is a look at Mistral Large 4’s page in the official documentation. Under “Other models,” you’ll find GLM-5.3 and GLM-5.2. The page showcasing the crown jewel of European sovereignty points, three lines down, to models built by Zhipu.
So both readings hold at once. Mistral can still build, and Mistral is already reselling what others build. That coexistence is precisely what deserves attention, rather than picking a side between the mourners and the cheerleaders.
The right benchmark isn’t Opus. It’s DeepSeek
Some were quick to compare Le Chonk with America’s flagship models and conclude that it lags behind. It does. Artificial Analysis’s Intelligence Index puts it at 38 points, between DeepSeek V4.1 Flash and GPT-6 Luna, OpenAI’s entry-level model. In a blind human evaluation of code quality, it came second out of five, well behind Claude Opus 5.
True, and beside the point. Mistral itself almost never measures itself against the Americans: its benchmarks pit it against DeepSeek V4 Pro, Kimi K3, GLM-5.3, and Qwen3.8 Max. Its stated goal is to be the best open model outside China, and on that front it delivers. The leap is striking, too: the same index gave Large 3 a score of 9 and Large 3.5 a score of 14.
An appealing theory is making the rounds on social media: Le Chonk is supposedly as good as anything the Americans produce with the same compute, so the gap now comes down to the size of the datacenters. It’s plausible. It isn’t proven: nobody publishes compute-matched comparisons, and Mistral makes no such claim. What has been established is more modest and more useful: Mistral now goes toe to toe with the Chinese labs, which also work under compute constraints.
And where most real-world use actually lives, in everyday office work, the gap is closing. On legal and financial tasks evaluated by Vals.ai, Mistral reports Le Chonk ahead of GPT-6 Astra. This is exactly what I called the good-enough theorem: the cutting edge stays American, but the overwhelming majority of use cases switch over as soon as a model is good enough at a much lower price.
Cybersecurity, or sovereignty through other people’s refusals
This is the strongest part of the announcement, and it’s subtler than most coverage suggests.
Le Chonk ranks in the global top five of the Artificial Analysis Cyber Index, an independent evaluation of how well models find and fix flaws in real software. On one of its tests, which asks a model to reproduce a vulnerability and then patch it, it scores 82%, the highest of any model.
But Mistral spells it out itself, and it’s the most revealing passage in the announcement: Claude Opus 5.5 and GPT-6 Astra score close to zero on that same test because they refuse the task. Part of the lead, then, reflects moderation policy rather than raw capability.
Far from weakening the argument, this is what makes it decisive. For a defender, proving that a flaw is real is often the first step toward fixing it. A model that refuses that work in the middle of a crisis, because a committee in California decided the request looked too offensive, is itself a vulnerability. Digital dependence isn’t measured only by the risk of having your access cut off. It’s also measured by moderation decisions made elsewhere, according to criteria you don’t control. Add downloadable weights and on-premise deployment, and you have exactly what the French state was after when it signed its framework agreement with the Armed Forces in January.
One question, however, remains unanswered. During the testing phase, cybersecurity experts, vetted partners, and public authorities get access to the same model with reduced moderation and expanded cyber capabilities. Who decides who makes the list, by what criteria, and with what audit trail? An offensive tool with discretionary access is only sovereign if it is controlled. Without a clear answer, we’ll simply have moved the moderation committee from San Francisco to Paris.
The balance of payments: the argument nobody puts numbers on
To my mind, this is the heart of the matter, and the reason this model matters more than its ranking.
We talk about AI sovereignty in terms of data (who reads it?) and dependence (who can pull the plug?). We forget the third dimension, the most prosaic one: who gets paid? Every subscription, every million tokens a French company buys from an American provider, is an import of services. And that import doesn’t just sit on top of existing spending: increasingly, it replaces payroll. When an accounting, legal, or marketing department shrinks its headcount thanks to an agent, the value that used to pay French employees now goes to servers and shareholders on the other side of the Atlantic.
The orders of magnitude are enough to show the imbalance. In May 2026, Anthropic’s annualized revenue had reached $47 billion. Mistral’s stood at around $400 million at the start of the year, with a target of $1 billion by the end of 2026. Part of that $47 billion comes from Europe, and therefore from France.
Yet for these uses, the race to the frontier isn’t endless. Progress at the cutting edge will remain decisive for scientific research and highly advanced applications. It no longer is for filling in a financial spreadsheet, drafting a legal memo, or triaging support tickets. On that ground, Le Chonk is already up to the job. The day a European model can run most of a company’s back-office functions, the hemorrhage becomes a choice rather than a technical inevitability.
Still, it’s worth stating what this argument doesn’t solve, or it slides into accounting chauvinism. The 3,800 GPUs are Nvidia chips, so part of the hardware spending flows back to California. The €3 billion Series D was led by Samsung, following a Series C led by the Dutch firm ASML: future dividends won’t all stay in France. The current promotional price is not a sustainable one. Above all, a French model only keeps money at home if people actually use it. Mistral claims more than 125 corporate customers in 20 countries and a partnership with Accenture, but a national champion left on the shelf has never rebalanced a trade deficit.
It’s the same logic I described regarding local AI, which you end up paying for three times: the heaviest bill isn’t always the one you see.
Deterrence through sufficiency
To understand what France can realistically aim for, the best analogy is a military one. Our nuclear deterrent never sought parity with the American or Russian arsenals. It sought a threshold: the point beyond which attacking us would cost more than an aggressor could hope to gain. Power isn’t measured by the size of the arsenal, but by the freedom to use it.
Applied to AI, the logic is the same. France doesn’t need the best model in the world. It needs a sufficient one, trained and served at home, that it can use without asking anyone’s permission: to defend its networks, run its administration and its industry, and stop its wealth from leaking out with every query. It’s the reasoning Dassault has followed for decades, and the one I defended regarding the SCAF fighter program.
But the analogy carries a requirement the enthusiasts tend to forget. A deterrent is only credible if the weapon is in your own hands. An arsenal stored with a host that also serves a rival’s models deters nothing unless you can take it back and run it yourself. Hence the importance of what is still to come.
What’s still missing
The license. The documentation lists Le Chonk as “Open” without publishing a single term. Large 3 shipped under Apache 2.0, Medium 3.5 under a modified MIT license: consistency is not guaranteed. Promising open weights three weeks ahead of release without saying on what terms is asking for trust that hasn’t yet been earned. Qobuz taught me that revocable ownership isn’t ownership. Sovereignty under a revocable license isn’t sovereignty either.
The harness. Value is shifting toward the orchestration and execution layer, the one that lets a model operate in complex environments and drive tools and agents in parallel. I made the case in my piece on Muse and accountability for AI agents: it’s the harness that makes the difference, not the model. On that front, Mistral has yet to show anything comparable to what the Americans offer.
Use cases. The Le Chonk announcement is an eleven-minute read packed with benchmarks. You learn what the model measures, much less what it lets you do that you couldn’t before. With OpenAI or Anthropic, every release comes with a concrete promise of what it’s for. With Mistral, you admire the feat. That’s the mark of a company whose first customers are the state and large corporations, clients to be reassured rather than won over. Yet stopping the hemorrhage described above will mean convincing companies that do have a choice.
My take
Le Chonk doesn’t catch up with anyone at the top, and it doesn’t need to. It reaches the threshold that matters, in the most sovereign of domains, trained on European machines. That’s a real achievement, and it would be absurd to sneer at it because it doesn’t beat Opus.
But my September 23 argument still stands: a model alone doesn’t make sovereignty. That will be decided on three points we can check over the coming months. A written, durable license when the weights are released. Special cybersecurity access governed by public rules. And above all, real adoption by French companies that stop sending their cash across the Atlantic.
Without that, we’ll have a national champion on the shelf, and a balance of payments that keeps bleeding.