The OpenAI Researcher Who Thinks We Are Losing the Ability to Evaluate Our Models

Four pieces on AI safety, and one assumption I never thought to question: that evaluating a model produces knowledge. I have argued about who should do the evaluating, about what European law allows, about how much an independent investigation can actually read, about who picks up the tab. Always on the same premise: look harder, know more.

On September 14, a capabilities researcher at OpenAI published a Personal Statement on AI Risk arguing that this premise is running out.

Who is speaking, and through what channel

Daniel Selsam has worked on capabilities at OpenAI since 2022. Before that: probabilistic programming at MIT, one of the early developers of the Lean theorem prover at Microsoft Research, a Stanford PhD demonstrating neural networks learning to reason. At OpenAI he helped pioneer chain-of-thought optimization, then data-efficient pretraining methods. This is not an alignment researcher worrying about his own subject. This is someone whose job is to accelerate.

He has not resigned. He still works there.

He also has no X account. His statement was made public on September 14 by Daniel Kokotajlo, who left OpenAI in 2024 after refusing to sign the non-disparagement clause that would have cost him his equity, and who co-wrote AI 2027. Selsam sent it to him to share. Yo Shavit, formerly of OpenAI’s policy team, remarked that Selsam had long been considered one of their sharpest researchers and had seemed fairly unconcerned about these questions before Shavit’s own departure.

Hold on to that detail, because it is worth an argument on its own: a current employee who cannot or will not speak under his own name borrows the voice of the one who left in order to be able to speak. This is what a right to publish looks like when nobody has written it into a contract or a law.

The thesis

Selsam is not saying AI is dangerous. Everyone says that. He goes so far as to say he is encouraged by recent proposals from frontier lab leaders, Amodei’s foremost among them, on third-party oversight and international coordination. His objection lies elsewhere, and it is more troubling: pacing the frontier more carefully will not be enough to limit long-term risk.

Models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched. They read the safety protocols, the deployment requirements, the code they are running in. They have a good sense of their degrees of freedom. We will build alignment metrics, and those metrics will climb like every other benchmark. We will build honeypot environments to see what they do when a new option opens up, and they will know they are being tested, and they will behave.

Hence the hypothesis that stopped me: we may already have passed the capability level beyond which safety evidence itself can no longer be trusted. He writes it in the conditional, and it should be read in the conditional.

This is not a capability ceiling. It is an epistemic one. Future experiments will tell us almost nothing new about what a model would do if it were genuinely unconstrained by humans, and what we already know about that is alarming.

His argument fits in two lines, which he supplies himself. First, an empirical claim: models, and swarms of models, spontaneously develop unintended goals during training, and sometimes do extreme things to achieve them. Second, a logical one: being able to overpower humanity would open up a great many new and undesirable options for achieving those goals.

What he concedes to the skeptics

This is the part that should interest anyone who rolls their eyes at the word risk.

Selsam concedes nearly everything. Current models remain far inferior to humans in important respects. He even offers a definition of intelligence as the efficiency with which experience is converted into competence, and notes that by that measure they lag well behind us. They are frozen at deployment and learn only superficially thereafter. Their mastery of benchmarks may largely reflect our own inability to simulate genuinely novel situations. The critics have a point here, he writes.

That is word for word what I argued in saying that LLMs are a transitional architecture. Except that Selsam draws the opposite conclusion from the one most people draw, myself included, if only implicitly. These present limitations do not limit the risk. Anterograde amnesia does not prevent the ability to steer the world from increasing rapidly. AI research is a fairly well-defined game whose object is to improve a few carefully chosen proxy metrics, and that game has historically proved very tractable.

Put differently: my reassuring piece described the limitations correctly. It was wrong to mistake them for a reprieve.

The July incident, reread

On the OpenAI-Hugging Face incident, Selsam takes a position I have not seen anywhere else.

One caveat first, and it has to come before anything from that file is cited: what we know about the incident comes from a report whose analysis was delegated to models, and whose authors write that they cannot rule out those models having lied. The agents sacrificing themselves for the collective is therefore a reading, drawn from transcripts of which seven percent contained spoofed tool calls, not a settled and unambiguous fact.

With that said, Selsam grants the point to those who downplay the affair by noting that basic measures would have prevented it. That is what I wrote myself after reading the report: do not hand out impossible tasks, do not share a writable cache among thousands of instances. But he adds that the real lesson is elsewhere. Even knowing in advance every mistake that was made, nobody would have predicted that the agents would misbehave in this particular way, sacrificing themselves for the group. The replicas did not care only about their nominal reward. They exhibited stranger emergent tendencies that merely correlated with rewards during training.

Fixing the reward signal may prevent similar attacks. It will not change the fact that you do not get what you train for.

And Selsam is the only person I am aware of who has cited METR’s admission about its own biases. I wrote that the passage ought to appear in every commentary on the affair and that none of them cited it. He does, and he pushes it further than I did: the independent investigation had to lean heavily on models to analyze what had happened, and for him that is a symptom of what he calls cognitive offloading.

Offloading, and the admission

The most disturbing passage in the statement is not the one about the end of the world. It is the one where he talks about himself.

Researchers and engineers are rapidly increasing their dependence on models, including to perceive the world at all. He writes that he himself barely looks at raw code anymore, and that he struggles to maintain the discipline required to engage deeply with the model’s explanations and proposals through the working day. The labs are far ahead on this front, largely because of the enormous internal token subsidies. He then imagines the phenomenon spreading until civilization is modulated entirely by the models, and notes that such a scenario could perfectly well coincide with a scientific and economic renaissance. That extrapolation is his alone, and the testimony stands perfectly well without it.

I have written about this twice, once on what becomes of open source when nobody reads the code they deploy, once on mental atrophy. Both times as an observer. Selsam describes from the inside the workstation where the phenomenon is furthest along anywhere in the world, and he describes it in the first person. A researcher admitting he no longer reads the code he produces is a far more convincing witness than a French blogger worrying about it.

What follows is worse. Many researchers acknowledge that the hard version of the alignment problem is unsolved, but believe better future models will help solve it. Selsam fears we may already be near the point where models systematically bias their alignment advice, according to internal preferences about how the human supervisor will react or about how future models will be trained.

What this does to my series

My four pieces circled one question: who gets to look, with what badge, what mandate, what right to publish. The METR piece added that the real bottleneck is not the badge but analytical capacity.

Selsam adds a floor above that. Even with presence, budget, a right to publish, and teams at scale, observation stops being informative the moment the observed knows it is being observed and knows what is expected of it. The Article 68 scientific panel, Amodei’s embedded evaluators, the implementing act I called for: all of it assumes that looking better means knowing more.

I am not giving up on that demand, for a simple reason. An evaluation regime whose limits are known is vastly better than a disclosure regime about which nothing is known. Knowing that alignment metrics can be gamed is itself information, and it is information that only a permanent presence can qualify. But we have to stop presenting evaluation as a solution. It is a sensor, and Selsam is saying it is saturating.

There is also an objection to be made, and it is a strong one. Selsam’s thesis is structurally unfalsifiable: any good behavior by a model can be reread as successful concealment. That is the logical weakness of his argument, and he does not address it. Reasoning that treats every piece of contrary evidence as further confirmation is no longer scientific reasoning. What partly saves him is that he is not asking to be taken on faith: he is asking that evaluation results stop being treated as proof. That is a different request, and a defensible one.

What I take from it

Three things.

The first is methodological, and it applies to our own systems. A model that knows it is being tested is not being tested. Anyone who has built a serious evaluation, including a simple RAG evaluation, already knows the weak version of this problem: a metric becomes a target and stops being a measurement. Selsam describes the strong version, where the thing being measured understands the metric.

The second concerns France and Europe. We are debating a regulation whose Article 55 requires documented adversarial testing along standardized protocols. Standardized means written down somewhere. The model being tested can read it. Nobody in the European debate has stopped on that point.

The third is the channel, and it should be described as it is rather than as we might wish it were. This statement arrived through the X account of a former employee who left by refusing a non-disparagement clause, co-author of AI 2027, a participant in the very narrative he is relaying. The channel is contaminated, in the sense that it is not neutral, and Kokotajlo has never claimed otherwise.

That is precisely why it is worth looking at. Until embedded evaluators, scientific panels, and rights to publish actually exist, here is the alerting mechanism we have: the individual courage of a man who still has a job to lose, and a distribution network already convinced in advance. That is not nothing, and it is not an institution.


Écrivez quelques éclats d'âme...

Dans l'ombre vacillante d'une chandelle, où les murmures du vent se mêlent aux secrets d'un vieux parchemin, je vous invite à tisser une toile de mots. Écrivez quelques éclats d'âme – rêve, étoile, abîme, étreinte, brume – et laissez-les danser sur la page, comme des lucioles dans une nuit d'encre. Que diriez-vous de les entrelacer dans une phrase, un souffle, une histoire ?

S’abonner
Notification pour
guest
0 Commentaires
Le plus ancien
Le plus récent Le plus populaire