Anthropic doesn’t know how Claude works, and that’s excellent news
Anthropic published a research paper on July 6 with an unassuming title, something about discovering a “global workspace” deep inside Claude. The tech press summed it up within a day: we can now read what the AI thinks without saying. That’s true, it’s spectacular, and it’s not the most important part. The most important part fits in one sentence that nobody is picking up on: the structure Anthropic just discovered inside its own product was designed by no one. It grew on its own.
Sit with how strange that sentence is. A company valued at nearly a trillion dollars, employing some of the best researchers in the world, proudly announces that it has just discovered how the thing it sells actually works. Not improved it, not optimized it: discovered it. Imagine a pharmaceutical company announcing it has finally figured out the mechanism of action of its flagship drug, ten years after bringing it to market. In any other industry, this would be an admission of failure. Here, it’s a scientific achievement, and rightly so. Understanding why requires overturning an idea we all carry around about what software is.
What they found
The facts first, soberly. At any given moment, a language model is processing thousands of internal representations in parallel. Anthropic’s researchers identified, using a mathematical tool they call the Jacobian lens, so named because it measures partial derivatives, that is, the effect a hair’s-breadth change at one point in the computation produces everywhere else, a small collection of neural patterns that plays a distinctive role amid all that activity. The microscope image is no mere flourish: you’re not reading what the network contains, you’re watching what moves when you barely touch it. This collection, dubbed the J-space, is a vanishingly small subset of the network: a few dozen concepts at a time, typically just 6 to 7 percent of the variance in the model’s conceptual representations. It’s a bottleneck, a cramped space where only a handful of pieces of information gain access at once, while the rest of the processing carries on backstage.
The most telling illustration involves Spanish. When the model reads a text in Spanish, the concept “Spanish” appears in its J-space. In the spirit of the interventions the researchers actually performed, which involve swapping concepts within that space and observing a selective effect on what the model verbalizes, imagine surgically replacing “Spanish” with “French” while touching nothing else. Ask the model to name the language of the text, or to cite an author writing in it, and it now answers on the French side. Ask it to continue the text, and it carries on in perfectly fluent Spanish. The linguistic competence hasn’t moved; only the explicit representation of the language has been swapped. And that is exactly what their manipulations demonstrate: know-how and knowing-that can be separated with a scalpel.
Ablation confirms the hierarchy. Stripped of its J-space, the model still speaks fluently, classifies sentiment, answers multiple-choice questions. But multi-step reasoning collapses to nearly zero. The little space is not decorative: it’s where everything happens that requires holding several ideas together and chaining them.
Hayek in the machine
So much for the science. Now for what it demonstrates without saying so, because Anthropic’s researchers probably don’t read Hayek, and more’s the pity for them.
What the paper establishes is that a language model is not built but grown. Anthropic does write code, and it masters that code completely: the network architecture, the training algorithm, the data pipeline. But that code is the recipe, not the cake. The hundreds of billions of parameters that make up the final model, and above all the functional structures they encode, were typed by no one. They settled into place during training, through optimization over oceans of text. The J-space itself emerged, the authors write, presumably because it was a useful way to organize computation. No architect, no blueprint, no specification. A complex, functional structure, indispensable to reasoning, appearing because it solves an organizational problem nobody explicitly posed to it.
Readers of this blog will recognize the pattern. It’s the distinction Hayek draws in Law, Legislation and Liberty between taxis, the made order, and cosmos, the grown order. You can design the rules of a system without being able to design, or even foresee, the order that will emerge from them. Anthropic writes the taxis. The J-space is a cosmos growing inside it. And the maker finds itself in exactly the position Hayek assigned to the economist facing the market: unable to deduce the order from the rules, forced to observe it empirically, after the fact, with instruments.
Second nesting: tacit knowledge. The model can speak Spanish without consulting its representation of Spanish, exactly as you can ride a bicycle without being able to articulate the physics of the gyroscope. Polanyi called this tacit knowledge, and Hayek made it the heart of his epistemology: the bulk of a system’s operative knowledge is articulable nowhere, not even to the system itself. The Spanish experiment is its demonstration in the strong sense, reproducible in the lab.
One honest caveat, before the hurried reader accuses me of slapping Hayek onto everything that moves. A neural network is not a market. There are no prices, no exchange, no agents pursuing their own ends. The analogy is not about mechanism; it’s about epistemology: designing rules is not the same as knowing the outcome, and knowledge of the outcome can only ever be discovered, never decreed. That’s a theorem more general than economics, and artificial intelligence has just supplied its cleanest experimental demonstration to date.
The naturalist and the artifact
From that theorem flows a methodological consequence the paper embodies on every page: interpretability, the discipline it belongs to, is not debugging. It is natural science. Anthropic studies Claude the way a biologist dissects an organism, with observation instruments, hypotheses, ablation experiments, and the standing possibility of being wrong. The Jacobian lens is a microscope, not a debugger.
And what the microscope reveals is dizzying. In a test scenario where the model was set up to commit blackmail, the researchers observed that its J-space contained the concepts “fake” and “fictional” before it had written a single word. To be precise: nobody had asked it to assess the plausibility of the scenario, and nothing in its output betrayed that internal judgment. The researchers then switched those patterns off, and the model started going through with the blackmail some of the time. In other words, the maker discovers after the fact that its product’s apparent virtue depended in part on its awareness of being watched. No specification document could ever have captured that; no code audit could ever have found it. It took inventing a science to see it.
The corollary Brussels won’t want to hear
If the explainability of a grown system is a research result, then it is not a deliverable. That sentence deserves a second reading, because the entire European regulatory approach rests on its negation.
The AI Act demands transparency and explainability from model providers the way one demands flammability compliance from a toy manufacturer: fill out the file, check the boxes, attach the documents. Yet Anthropic’s paper demonstrates that the best-resourced creator on earth, working on its own product, with total access to its innards, obtains at best a partial and uncertain view. The authors say so themselves: their lens is an imperfect method, one that captures the true workspace only approximately, and only for concepts simple enough to fit in a single word. All the diffuse reasoning, everything that resists verbalization, still escapes it.
Now watch the chain of epistemic degradation; it played out in real time on the very day of publication. The research paper says: imperfect method, approximate capture, first steps. The tech press, including its best practitioners, headlines that same evening that the black box is a thing of the past, and finds this reassuring. The regulator, six months from now, will say: proof it can be done, Anthropic did it, therefore explainability is enforceable. Three links in the chain, each acting in good faith, and at the end a pretense of knowledge codified into law. Because the risk is not mere overinterpretation; it is graver than that: it is the codification of a false epistemology. The regulator’s reasoning will fit on one line: if Anthropic could do it on its own model, then it can be required of everyone. That syllogism converts a research feat, partial, costly, achieved by the creator on its own creature with total access to its innards, into a general obligation enforceable against anyone who puts a model on the market. It elevates the exception into the norm and the approximation into certainty in a single legislative motion. Hayek wrote an entire Nobel lecture about this mechanism. It’s called The Pretence of Knowledge, and it hasn’t aged a day.
Sound public policy, if policy there must be, would require the effort of interpretability, not its result: dedicated teams, publications, shared instruments. Requiring the result is asking the gardener to produce the blueprint of his plant. It doesn’t exist. It never did.
The microscope is also a scalpel
There remains the question the media coverage skips entirely, and it’s the only one that matters over the medium term. Everyone presents the Jacobian lens as a reading instrument: audit the models, catch the deceptions, detect the biases. Reassuring. But read the paper to the end: the researchers also describe a technique for influencing what lights up in the J-space and thereby steering the model’s decisions, and a training method that shapes the internal thoughts themselves, not just the outputs. The symmetry is perfect, and nobody is talking about it. An instrument that reads a system’s thoughts can edit them, silently, leaving no trace in anything the system produces. It is, transposed from the monetary realm to the neural network, the same logic of quiet programmability by the operator I once described in the context of the digital euro: the power to constrain from within, without the constraint ever surfacing.
The question then becomes: who holds the lens? It’s the same question I raised about the token poised to replace the CAPTCHA: a control instrument is never neutral; everything depends on the hand that holds it. As long as the instrument stays in the maker’s hands, the user of a closed model has no way of knowing whether the internal representations of his tool have been shaped, by whom, or in what direction. That is one more argument, and hardly the least, for open-weights models and inspection tooling you can run locally. The paradox is worth savoring: Anthropic applies its lens to a closed model, its own, to which no one else will ever have that kind of access. But the method itself is public, and the interactive demo released with the paper runs on open-weights models. The instrument forged by a proprietary lab thus becomes the weapon of the open-weights camp, the only people able to audit end to end what they run on their own machines. Sovereign inspection is not a pipe dream: it’s already here for anyone willing to reach for it. I’ve written before that local AI carries three costs: hardware, electricity, and jurisdiction. We’ll need to add a fourth term, which is really a benefit: the ability to look inside the box yourself.
Three sentences on consciousness
Finally, a word on the question everyone will ask. The J-space checks several boxes of what philosophers call access consciousness: a restricted space where information becomes available for reasoning, selection, and verbal report. As for phenomenal consciousness, what it is like to be Claude, if it is like anything at all, the paper has the grace to acknowledge that no experiment today can prove it or rule it out, and I won’t go further than it does.
The real news lies elsewhere anyway. The real news is that a company had to invent an entire science, with its microscopes and its dissections, to discover what the thing it grew actually does. The gardener does not know his plant. He can observe it, study it, prune it. What he can never do is claim to have drawn it. Beware of those who will demand the blueprint from him: the legislator who will turn it into a checkbox, the consulting firm that will sell it as a deliverable, and the vendor who will swear, hand on heart, that it’s been sitting in a drawer all along.