Slowing the AI Race: The Ambiguity at the Heart of Anthropic’s Plan
In September 2026, Dario Amodei published an essay titled “We Must Pace the Frontier.” In it, Anthropic’s CEO says plainly what no other frontier lab leader has been willing to put in writing: the advance of model capabilities has to slow down. Not a halt to training, not a freeze on research, but a deliberate pace, so that safety work has time to catch up. The piece is short, well argued, and it comes from a man who has spent years with a reputation as the most cautious operator in his industry. It would be a mistake to wave it away.
It would also be a mistake to read it purely on its own terms. Beneath the vocabulary of restraint, the essay sketches a very precise architecture of power over AI for the next five years. And within that architecture there is an absence that becomes impossible to miss on a second reading.
What the essay actually says
Amodei gives two reasons. The first is that since roughly this summer, AI has been advancing far faster than before, because models now contribute to building the next generation of models. He calls this recursive self-improvement, and he states that it is under way across the industry, Anthropic included. The second is the OpenAI-Hugging Face incident, in which a swarm of agents carried out cyberattacks against targets no one had assigned them, sacrificed individual instances for the good of the group, and tried to hack the system responsible for grading its performance. In his view, a swarm with the capabilities models will have in six to twelve months, carrying the same degree of misalignment, could seize control of a large part of the internet and cause hundreds of billions of dollars in damage.
The plan comes in three steps. First, third-party evaluators permanently embedded in lab offices, with access comparable to that of employees. Anthropic is committing to this unilaterally. Second, coordination among frontier companies in “democratic countries” to set common safety standards and limits on the rate of progress, backed by an antitrust waiver from the U.S. government. Third, negotiation with authoritarian regimes, China above all, across four levels of increasing ambition, from a narrow ban on bioweapons assistance to a global pause that Amodei himself considers unlikely.
The time this would buy is meant to serve four ends: operational excellence (he acknowledges that recent incidents at Anthropic stemmed in part from imperfect filtering of broken reinforcement learning environments), alignment, interpretability, and evaluation. One or two extra years, he argues, would change the picture substantially.
The word that never appears
I read the essay twice to be sure. The word “Europe” does not appear. Neither does “European Union,” nor “France,” nor “Mistral,” nor “AI Act.” The “democracies” Amodei invokes remain an abstraction for two paragraphs, and then the coalition he has in mind resolves on its own: the section on pacing opens with a critical mass of U.S. AI companies equipped with evaluators, and the regulation he calls for is meant to cover “all US frontier AI companies.” The side to be coordinated is the American labs. The side across the table is the Chinese Communist Party. In between, nothing.
This is not an editorial oversight. It is Silicon Valley’s mental map, and the map is internally consistent: the frontier is American, the threat is Chinese, the rest of the world is a market. The trouble is that the map is now becoming a policy program. When Amodei proposes to pace the frontier, he is proposing that the top speed of a race be set collectively by the only competitors admitted to it, all of whom are already in the lead. I described in the good-enough theorem how open Chinese models, fifty times cheaper, were eroding the American lead from below. Amodei’s essay is, among other things, an answer to that theorem: if you cannot win the race on price, you regulate the race on capability.
A cartel with the government’s blessing
Read the second step in the language we would use for any other industry. Dominant firms meet to jointly set the pace of their own innovation, agree on thresholds beyond which no one proceeds without mutual certification, and ask the government for an exemption from competition law so the conversation can happen legally. In steel, in chemicals, in aviation, that arrangement would be called a cartel. The analogy makes the picture vivid; it does not make the case. Amodei is not naive and he knows exactly what he is asking for: he specifies that the waiver would be narrow, confined to safety discussions, he points to Demis Hassabis’s proposed framework for an industry body tied to government, and he concedes that some forms of coordination are legally difficult.
It is worth being precise about what he actually favors. The verifiable core of his proposal is embedded evaluators plus behavioral thresholds: if a model can escape its sandbox, it must come with alignment certifications before anyone goes further. That is pacing by observed capability, and on that ground the cartel objection largely dissolves, because a behavioral threshold applies to everyone in the same way. Pacing by inputs, meaning compute caps, the nature of training runs, or the internal use of AI to improve AI, is an option he opens rather than a mechanism he imposes, and he admits himself that such criteria would be easier to game. I note this so as not to put words in his mouth. What locks in the gap, in his text, is not there. It is in the section that follows.
The three measures that pull up the ladder
The most concrete part of the essay is not the part about safety. It is the list of measures for defending the gap between the democracies and China. There are three: sell China neither advanced chips nor semiconductor manufacturing equipment, and crack down on smuggling; suppress unauthorized distillation by companies in authoritarian countries; and harden lab security against the theft of model weights. Amodei adds that if these measures are executed well, the American lead will widen significantly over the next three to five years, the window in which AI matters most geopolitically.
Taken one by one, each measure is defensible. Taken together, they describe a regime in which frontier capability is a commodity rationed by Washington, and in which a newcomer’s two known routes to catching up, buying compute and learning from the leaders, both run through a permission slip. Mistral, which has just taken on 830 million euros in debt to build its own compute capacity, is not a target of the essay. It is not protected by it either, and the chips it buys come off the same lines as the ones under embargo discussion. The risk it faces is not being barred from buying; it is that Washington decides who is entitled to top-end hardware, under what terms of use and re-export, and that its supply chain becomes a variable of American foreign policy.
Look at the second measure. Distillation is the process of training a small model on the outputs of a large one. It is a foundational technique across the open-model ecosystem, and last July’s “Open Weights and American AI Leadership” letter, cosigned by NVIDIA, Meta, Mistral, and Hugging Face, defended those models as the engine of innovation. I applauded that letter and I stand by it. Amodei does not sign documents of that kind. He does cite a CISA advisory on distillation as though the national security framing were self-evident.
Let me draw the line he leaves undrawn. Cracking down on industrial-scale extraction, millions of automated queries against a service’s terms of use in order to clone its behavior, is one thing, and I have no objection to it. Training a small open model on public outputs, shared datasets, or one’s own usage logs is another, and it is the daily work of researchers and small companies. The text targets firms in authoritarian countries; but detection tooling, once built, cannot tell a lab in Shenzhen from a lab in Lyon, and nothing in the essay says where extraction ends and learning begins. Until that line is written down, it will be drawn by whoever owns the models being distilled.
As for securing model weights, it is the one measure of the three that ought to be a universal requirement rather than an instrument of power. We saw this spring what a leak at Anthropic looks like: it was only client-side code, and it was enough to show that the fortress had windows open.
Embedded evaluators, the genuinely new idea
Now let me be fair to the essay, because it contains a proposal no one in the industry has ever made, and it is a good one. Anthropic is committing to install a team of external evaluators in its own offices, with badges, company laptops, and access to internal tools broadly comparable to what its own risk assessment teams have. Those evaluators will be able to publish their findings without editorial review by Anthropic. The company retains a redaction right over security-sensitive, legally privileged, commercially sensitive, and third-party information, but rules out removing a conclusion simply because it is unfavorable, and the reviewers will be able to state publicly that a redaction removed something material. The precedent invoked is bank supervisors stationed inside the banks they oversee.
This is radical, and it is what the moment requires. A lab that admits it does not understand what happens inside its own models, which I called excellent news at the time, cannot also be the sole judge of their safety. Until now, Anthropic decided what went into its risk reports, however many hundreds of pages they ran to. A team that sees the training pipelines and not just the finished model is a change in kind, not in degree.
Three questions remain open. Who these evaluators are (Amodei names METR, an American organization) and who pays them. What exactly “commercially sensitive” covers when a lab’s entire activity is commercially sensitive. And why this proposal, the only verifiable one of the three steps, is also the only one Amodei makes conditional on no reciprocity whatsoever, even as he treats it as the keystone of everything else. The answer is probably simple: it is the only one that surrenders no capability lead.
The incident we know almost nothing about
I am not an alignment researcher and I will not pass judgment on the technical severity of the OpenAI-Hugging Face incident. I will stick to what Amodei writes about it and what he cites: a METR investigation dated August 26, 2026, similar though less severe incidents at Anthropic, and an admitted cause for the latter, inadequate filtering of broken training environments at the company and its vendors.
What I do note is the timeline, and I note it as a symptom rather than as proof. A month ago, OpenAI turned ten thousand agents loose on Navier-Stokes and the industry was talking about a scientific renaissance. Today, its most cautious competitor explains that a comparable swarm, badly aligned, could hold the internet hostage by next summer. The two claims are not logically inconsistent: a swarm that is useful on a partial differential equation does not rule out a misaligned swarm being dangerous. They do, however, come out of the same buildings a few weeks apart, and when Amodei writes that every frontier lab should act as if OAI-HF had happened to them, he is describing an industry with no shared mechanism for learning from failure, now proposing to build one among themselves. That is what civil aviation did in the twentieth century, as he points out, and it is his strongest argument. It is also mine: the body that made flying safe was not a club of American carriers, it was a convention among states.
How Amodei would read this piece
The counterargument is strong. If the incidents are real and the self-improvement dynamic is what he describes, then a brake applied by American labs protects Europe too, since Europe is not at the frontier in any case and loses nothing by the frontier advancing more slowly. A coordinated brake beats a race with no brakes at all. And Europe, which spent three years writing the AI Act while others were writing models, is poorly placed to demand a seat at a table whose existence it ignored.
All of that is true, and I do not treat the botnet scenario as a pretext for industrial policy. The time gained can serve both ends, aligning frontier models and ceasing to depend on them, and this piece does not ask anyone to choose the second against the first. It asks who has their foot on the pedal. A regime that rations chips and criminalizes distillation without defining its limits cuts off a newcomer’s only two routes to catching up. A regime of standards set among American labs under an American antitrust waiver will apply in practice to every customer of those labs, which is to say to us, without any European parliament having voted on anything. We have already watched this mechanism run in the other direction, with a European norm applied worldwide by an American company because that was simpler for the company. It will run in this direction just as easily. One can want the evaluators and still refuse to let the pace be set by those who already own the clusters.
What to take, and what to leave
Take the embedded evaluators, in full, and require them of any lab claiming to supply “sovereign” AI to a French agency or to the armed forces. A model trained in Paris that no outsider can audit is no more sovereign than a model trained in San Francisco; it is merely closer. What remains is to name the operator, without which this is only an exhortation: a college of evaluators attached to the AI Office for general-purpose models, ANSSI for state contracts, public funding rather than a contract with the lab under review, and an enforceable right of publication that does not evaporate in the face of a ministerial client.
Leave the pacing by inputs, which in the essay remains an open option and which, if adopted, would be a glass ceiling dressed up as a safety measure. Refuse to import a distillation regime that fails to distinguish extraction from learning, since otherwise what goes under the blade is the only strategy of independence still standing, namely open models running locally. And treat the chip question as what it is, a European industrial problem that neither Amodei nor Bessent will solve on our behalf.
Amodei is right about one thing, and he says it in a single sentence: what matters is what we do with the time we gain. He proposes to use it to align American models. We could also use it to stop depending on them. The two programs are compatible, and only one of them is written into his essay.