AI text detectors : the tool that should never have existed

Universities are punishing students on the strength of a score produced by an algorithm its own creators warn against using as proof. Professors brandish percentages as if they were confessions. Theses are voided. Careers damaged.

This isn’t a problem of a poorly calibrated threshold. It isn’t a problem of a model that simply isn’t good enough yet. It’s an error of principle, and it rests on a complete misunderstanding of what a language model actually is.

What Nobody Says Out Loud in Disciplinary Hearings

An AI text detector is a binary classifier. Its job: given a text T, decide whether it came from a human or a machine. On the face of it, the problem seems reasonable. In practice, it’s intractable.

For a binary classifier to work, the two classes it’s trying to separate have to be statistically distinct. That’s the baseline condition. Not sufficient, but necessary.

And large language models are trained precisely so that this condition is never met.

KL Divergence: The Objective That Makes Detection Impossible

When you train a large language model, the goal is to minimize the Kullback-Leibler divergence between the model’s distribution and that of human text:

D_KL(P_human ‖ P_AI) → 0

In plain, non-mathematical terms: you train the model to produce text whose statistical distribution is identical to that of a competent human. That training, incidentally, rests on a vast corpus of human writing, and largely without consent, as I detailed in LLMs: The Ultimate Copyright-Laundering Machine. That is the very definition of success. A perfect LLM would produce text that’s indistinguishable: not to the naked eye, and not by any statistical method either, because the distributions would be, by construction, the same.

An AI text detector is therefore trying to separate two distributions whose explicit training objective is to converge. It’s structurally absurd, and it only gets worse as the models improve. I’ve already explored the internal geometry of LLMs: what we take for intelligence or consciousness is nothing more than vector distance. Statistical detection runs into exactly the same wall.

Antoun, Mouilleron, Sagot, and Seddah documented this rigorously in a study presented at JEP-TALN-RÉCITAL 2023, the leading French-language conference in natural language processing: existing detectors perform decently within their training domain but collapse the moment you change style or context. They’re vulnerable to the most trivial attacks, heavily dependent on a didactic register, and their conclusion is explicit: extreme caution for any real-world application. In institutional terms: unusable for sanctioning anyone.

Watermarking: The Only Serious Avenue, and Why It Changes Nothing for Universities

Defenders of detection will have one argument in reserve: what about watermarking? Approaches like Google DeepMind’s SynthID don’t rely on surface statistics. They embed a cryptographic signal directly into the generation process: a kind of invisible signature in the token-selection probabilities. It’s a radically different idea, and an intellectually honest one.

Except that Google itself documents the central limitation: the signal collapses the moment you reword the text. A round-trip translation, a paraphrase, and detection falls to levels close to random noise. Researchers at ETH Zurich showed in 2024 that unsophisticated adversaries can strip SynthID’s watermark with success rates above 90%.

But the fundamental problem lies elsewhere: watermarking only works if the text hasn’t been altered, and only for the models that have integrated the system. A student who uses a third-party LLM, or who simply rewrites the generated text, is completely invisible. Watermarking is a solution to the problem of provenance at scale, not a tool for classroom surveillance. Using it as one would mean condemning the absence of a signature rather than the presence of fraud.

Post hoc detection remains a fantasy. Statistical or cryptographic, it makes no difference.

What These Tools Actually Measure, and Why That’s a Problem

The two metrics that dominate these detectors:

Perplexity. A low-perplexity text is a predictable one, in which the tokens follow one another along high-probability paths. The idea: LLMs supposedly write text that’s “too smooth.” Except that a lawyer, an engineer, a biologist, anyone writing in a codified technical register, produces exactly the same profile. Precision and clarity are academic virtues. On these tools, they count against you.

Burstiness. Humans supposedly have variable perplexity: complex sentences mixed with short ones, shifts in register. LLMs are said to be more uniform. Another heuristic that falls apart the moment you leave informal text behind, and one that systematically penalizes non-native authors, whose writing is naturally more homogeneous.

In both cases, what’s being measured isn’t the origin of the text. It’s the style. And punishing a style is precisely what an academic institution should never do.

The False Positive Has a Name: A Baseless Accusation of Cheating

In a binary classifier operating on overlapping distributions, the false-positive rate isn’t a technical detail: it’s the heart of the problem.

In an academic setting, a false positive doesn’t stay an abstract statistical error. It becomes a summons. A disciplinary file. Sometimes an expulsion.

The most exposed victims are predictable: foreign students whose academic writing is more formal and less variable than a native speaker’s. Students in scientific or legal programs whose register is codified by its very nature. Good students, quite simply: the ones who write with precision and clarity.

Documented cases exist in the United States, the United Kingdom, and Australia: theses annulled, degree conferrals put on hold, exhausting appeals, and in several cases, decisions ultimately reversed after human review, sometimes only after months.

Universities Sanction Students for Alleged AI Use, Now They're Suing -- Minnesota & Yale Cases

Universities like Vanderbilt disabled Turnitin’s detector back in 2023, after internal testing, noting in passing that even at an advertised false-positive rate of 1%, that amounted to 750 potentially erroneous cases out of the 75,000 papers submitted each year. In Australia, the Australian Catholic University ultimately suspended the tool after reports of students stuck for months in proceedings based solely on an algorithmic score. GPTZero posts its disclaimers in fine print. No one on the disciplinary committees reads them.

On July 20, 2023, OpenAI withdrew its own AI text detector (six months after launch), publicly admitting a true-detection rate of just 26%, with 9% false positives on human-written text. The arms maker admits its own radar is blind. The generals keep buying Turnitin licenses.

ChatGPT: Academic misconduct & AI content generator & detector

And for anyone who still doubts how robust these systems are: studies from 2024 and 2025 show that a simple paraphrase, or a prompt as trivial as “write like a non-native student,” drops the detection rate below 20%. So the tool offers no protection against deliberate cheating: it only punishes those who didn’t think to get around it.

The Institution That Mistakes Probability for Proof

What strikes me in all this isn’t the mediocrity of the technology. It’s the intellectual carelessness of those who adopt it.

Institutions that claim the mantle of scientific rigor use a probabilistic tool as if it delivered binary certainties. A score of “87% AI” is not proof. It’s an estimate (over overlapping distributions) produced by a system whose own makers won’t vouch for its reliability in a disciplinary context.

No court would accept such evidence. No research ethics committee would sign off on a methodology with that error rate. Academic panels, however, will.

Under French university disciplinary law, a sanction must rest on materially established facts and respect the adversarial principle, the student’s right to confront and contest the evidence. Yet an AI detector is by definition a black box: Turnitin doesn’t publish its methodology, doesn’t document its thresholds, and its scores shift from one version to the next. How is a student supposed to challenge a score whose origin and computation they don’t even know? A sanction resting on that single element is legally fragile, and an appeal before the CNESER or the Conseil d’État on those grounds would stand a good chance of succeeding.

This is statistical illiteracy, institutionalized. And it has real victims.

What Universities Should Do

The real answer to AI in assessment settings isn’t to hunt it down: it’s to rethink how we assess. Oral exams, in-person work, process tracking, research journals, thesis defenses. Pedagogical answers to a pedagogical shift.

UCT drops AI detection tools in shift toward ethical tech use

The algorithmic witch hunt protects nothing. It strikes at random, punishes those who write well, and gives a clear conscience to institutions that never wanted to do the work of adapting.

An AI text detector isn’t a tool in the process of being improved. It’s a tool whose very principle contradicts the mechanics of the models it claims to detect. Deploying it to sanction a student means condemning someone on an unfavorable statistical draw.

In any other discipline, you’d call that a methodological error. In a disciplinary panel, they call it a sanction.


Écrivez quelques éclats d'âme...

Dans l'ombre vacillante d'une chandelle, où les murmures du vent se mêlent aux secrets d'un vieux parchemin, je vous invite à tisser une toile de mots. Écrivez quelques éclats d'âme – rêve, étoile, abîme, étreinte, brume – et laissez-les danser sur la page, comme des lucioles dans une nuit d'encre. Que diriez-vous de les entrelacer dans une phrase, un souffle, une histoire ?

Subscribe
Notify of
guest
0 Commentaires
Oldest
Newest Most Voted