Skip to content
🤖 Artificial Intelligence

The detector is scoring your English, not your honesty

Francis Okafor Francis Okafor
10 min read
ai-detection academic-integrity watermarking non-native-english higher-education china
The detector is scoring your English, not your honesty
On this page
  1. Seven detectors, ninety-one TOEFL essays
  2. Perplexity is a fluency meter
  3. A watermark records involvement and nothing else
  4. Two yuan per thousand characters
  5. The universities have a real problem
  6. What survives a denial
  7. The result nobody quotes
  8. Sources

《滕王阁序》 was written around 675 AD by Wang Bo, a Tang poet who drowned in his twenties. A widely used Chinese AIGC checker scored it at close to 100 percent machine-generated.

Zhu Ziqing's 《荷塘月色》, an essay from 1927 that most Chinese schoolchildren still memorise, came back at 62.88 percent on one commercial detector. A different Chinese checker returned 0.0 percent on exactly the same text.

Nobody circulating those screenshots names the detecting system consistently, and I cannot reproduce either score. The unreliability is the finding, not a caveat on it. Two systems given identical prose from 1927 produced numbers 62.88 points apart, and both presented the result with the confidence of a lab assay.

Seven detectors, ninety-one TOEFL essays

The paper everyone cites badly is Liang, Yuksekgonul, Mao, Wu and Zou, published in Patterns in July 2023. They took 91 TOEFL essays written by non-native English speakers, pulled from a Chinese educational forum, and 88 essays by American eighth-graders from the Hewlett Foundation's Automated Student Assessment Prize dataset. Then they ran both sets through seven detectors: Originality.AI, Quil.org, Sapling, OpenAI's own classifier, Crossplag, GPTZero and ZeroGPT.

Average false positive rate on the TOEFL essays: 61.22 percent. Eighteen of the 91 were flagged as AI-generated by all seven detectors simultaneously. Eighty-nine of the 91 were flagged by at least one. On the thirteen-year-old Americans the same seven tools were close to perfect.

Then the authors ran the experiment that should have ended the argument. They fed the human-written TOEFL essays back through ChatGPT with a single instruction, to enhance the word choices to sound more like a native speaker, and re-scored them. The false positive rate fell to 11.77 percent. Exactly one essay in 91 was still unanimously flagged.

Ninety-one essays is a small sample and the authors say so. The detectors were 2023 vintage, most of them perplexity scoring against a GPT-2 backbone, and OpenAI quietly withdrew its own classifier ten days after the paper appeared. None of that rescues the mechanism, because the mechanism is still what everything is built on.

Two Chinese checkers, given the same 1927 essay, returned 62.88 percent and 0.0 percent, and both printed the number to two decimal places.

Perplexity is a fluency meter

Here is what a perplexity-based detector actually computes. It takes your text, feeds it to a reference language model and asks how surprised that model is by each word given the words in front of it. Low surprise across a whole document reads as machine. Burstiness is the companion metric, the variance in that surprise from sentence to sentence. Humans are supposed to spike. Models are supposed to flatline.

Both quantities are properties of the finished text. Neither is a property of the process that produced it. A person can write flat prose by hand and a model can be prompted into jagged prose, and the detector has no channel through which to learn which of those happened.

I have lived in Shenzhen for eight years and I work in Mandarin every day. My Chinese is fluent and it is also conservative. Messaging a supplier in Longhua about a delivery slipping, I will write 没问题 a hundred times before I reach for 不成问题. Not because the second phrase is unfamiliar to me. Because the first one has never once been misread, and a misread message on a Thursday afternoon costs a day of production.

That instinct does not switch off when the language changes. A Nigerian engineer writing a personal statement in English, a Vietnamese postgraduate writing a methods section, both reach for the construction they are certain of and then reuse it. Narrow vocabulary band. Predictable collocations. Low variance across sentences. Which is, line for line, the signature a perplexity detector was built to catch.

Nobody is measuring dishonesty here. The quantity on the screen is how far your English sits from the middle of a distribution assembled mostly out of native prose.

A watermark records involvement and nothing else

Watermarking gets offered as the grown-up answer, and it is genuinely a different technique. SynthID-Text, published by Google DeepMind in Nature in October 2024, does not analyse finished text at all. It is a logits processor sitting inside the generation pipeline after Top-K and Top-P sampling, biasing the model's token probabilities using a pseudorandom function keyed on the preceding words. Detection is a Bayesian score against a threshold, returning one of three states: watermarked, not watermarked or uncertain.

Google's own developer documentation lists the limits without much spin. Watermark application is less effective on factual responses, because there is less room to shift word choice without damaging accuracy. Detector confidence drops sharply when the text is thoroughly rewritten or translated into another language. And the detector reads SynthID and only SynthID.

Anthropic shipped a version of the same scheme on Claude on 14 August 2026, under the EU Code of Practice on Transparency of AI-Generated Content. Their explanation is the most useful document currently available on this subject, because it states the limit in the first person. The watermark can establish that Claude was likely involved with the content at some point. It cannot distinguish Claude wrote this from Claude heavily edited this. Involvement. Not authorship, not proportion, not contribution.

The negative case is weaker still. Unwatermarked text might be human. It might come from a model that does not watermark. It might be watermarked output that somebody rewrote past the signal, or a passage short enough that the signal never accumulated. A clean result is not evidence of a human hand. It is an absence of evidence about one specific model.

Two yuan per thousand characters

CNKI, the state-linked academic database most Chinese universities push their theses through, added AIGC detection in 2024 at roughly two yuan per thousand characters. By the 2025 graduation season more than a dozen institutions including Fuzhou University, Sichuan University and Jiangsu University had capped suspected AI content somewhere between 15 and 40 percent. Heilongjiang University published a scale: 40 to 70 percent triggers a six-month revision period, above 70 percent the degree qualification goes.

A market appeared within weeks, which is what markets do here. Rest of World documented students paying around 500 yuan for a tutor to rewrite a thesis by hand, 70 yuan for access to a pre-submission testing platform and as little as 16 yuan for a startup tool that scrambles sentences until the number drops. A German literature senior in the northeast, writing under the name Xiaobing, watched a 16-page paper she had written herself come back at 50 percent AI.

One blogger paid 120 yuan to run a 58,000-character thesis he wrote by hand through CNKI's checker. It returned 86.8 percent.

I watch this from a specific vantage point. Everything above is happening in Chinese, to Chinese students, using detectors tuned on Chinese academic prose, and the failure mode is identical to the one Liang measured in English. Formal register. Precise terminology. Long subordinate clauses. Anything that reads as disciplined academic writing pushes the number up, which is why the reported cases keep landing on people who wrote carefully.

Students on both sides of that corridor have worked out the same response, independently, in two languages, without reading a single paper about it. Write worse on purpose.

The universities have a real problem

The honest counter to all of this is that institutions are not being stupid. They are cornered. Turnitin marked the first year of its AI writing detector in April 2024 with a number: over 200 million papers reviewed, of which more than 22 million, roughly 11 percent, contained at least 20 percent AI writing, and over six million, roughly 3 percent, contained 80 percent or more. Those are vendor figures and should be read as vendor figures. They are also not nothing.

Human judgement performs worse. The University of Sydney's teaching group pulled the relevant studies together in October 2025: experienced educators correctly identified AI-generated text 38 percent of the time in one study, and in a live examination setting, 94 percent of AI submissions went undetected by human markers.

So the claim that detectors are biased is true and, standing alone, useless. Turn the detector off and fairness has not been restored. What has been restored is a system where a marker's private hunch about who sounds like they wrote this operates entirely unchecked, and hunches of that kind have a well-documented direction of travel.

Newer commercial detectors also do genuinely better on second-language writing than the 2023 cohort managed. Several vendors now publish near-zero false positive rates on ESL corpora and some independent 2026 evaluations broadly agree with them. I am not quoting those figures as established, because most of the impressive ones are first-party, and a false positive rate a company measured on itself is a marketing claim wearing a lab coat.

What survives every improvement is the base rate. Vanderbilt did the arithmetic in public on 16 August 2023 when it disabled Turnitin's detector: 75,000 papers submitted in a year, a vendor-claimed 1 percent false positive rate, roughly 750 students wrongly flagged annually at one university. Cut the error rate tenfold and you still manufacture 75 accusations a year against people who did nothing. A March 2026 arXiv paper by Nathan Garland formalises why this does not engineer away: any text-only, single-shot detector with real statistical power has to produce false accusations at a rate governed by how far student writing and model output overlap. That overlap is not a defect. It is what you get when models are trained on human writing and humans are trained to write in registers.

What survives a denial

Turnitin's own guidance, published in June 2023, tells assessors to treat the highlighted sentences as a reason to initiate a conversation, not to draw a conclusion. Almost nobody does. The number appears, the student is summoned, and the conversation is over before it starts. What actually holds up when a student says I wrote that is process evidence, and it holds up because it is expensive to fake in the specific way that matters.

Version history first. Timestamped revisions in a document, git commits, an edit log containing forty hours of small ugly changes. Sydney's guidance concedes the obvious objection, that revision histories can be manufactured. True. Manufacturing a plausible six-week edit trail is a different and much larger act than pasting an answer, and it moves the dispute from a statistic onto a piece of conduct that can actually be adjudicated.

Oral defence next. Fifteen minutes, in person, asking someone to explain a decision inside their own paper and then to defend a choice they did not make. The person who wrote it manages this without effort. The person who did not fails visibly, in front of everyone in the room including themselves. It is the only method here that both works and scales badly, which is precisely why it stays out of the policy documents.

Artefacts around the work carry weight too. A discarded outline. An annotated bibliography that is wrong in an interesting way. A supervision note from six weeks before submission. And disclosure has to be safe to make: if writing I used Claude to tighten section three carries no penalty, most people write it, and if it carries a penalty the disclosure vanishes and you are back to guessing from perplexity scores.

None of this is free. Oral defence at scale costs staff hours nobody has budgeted for, and a detector costs two yuan per thousand characters. That is the real argument, and it deserves to be had in those terms rather than dressed up as a question about accuracy.

The result nobody quotes

Go back to the second experiment in the Stanford paper, because it is the part that never survives into the coverage.

Ninety-one essays written by human beings. Sixty-one percent flagged as machine-generated. Feed those same human essays through ChatGPT with an instruction to sound more like a native speaker, and the flag rate falls to twelve.

The student who launders their honest work through a language model passes. The student who submits it exactly as they wrote it gets called in.

I have sat with young engineers in Lagos and in Shenzhen who reached that arithmetic on their own and adjusted accordingly. They are not cheating. They are optimising against a measurement, which is what people have always done to a measurement. Somewhere inside that loop is a student who wrote every word herself, refused to launder it and lost the place.

Sources

Liang, Yuksekgonul, Mao, Wu & Zou, GPT detectors are biased against non-native English writers (Patterns, July 2023) : https://arxiv.org/abs/2304.02819

Anthropic, How Claude's text watermarking works (14 August 2026) : https://www.anthropic.com/news/claude-text-watermark

Google, SynthID: tools for watermarking and detecting LLM-generated text (developer documentation) : https://ai.google.dev/responsible/docs/safeguards/synthid

Rest of World, Chinese students are using AI to beat AI detectors (2025) : https://restofworld.org/2025/ai-detector-software-workaround/

163.com, Chinese report on AIGC detection rates for Zhu Ziqing, Wang Bo and a hand-written thesis : https://www.163.com/dy/article/KTI3TDSH0525CHJG.html

Turnitin, One year anniversary of its AI writing detector (April 2024) : https://www.turnitin.com/press/turnitin-first-anniversary-ai-writing-detector

Vanderbilt University, Guidance on AI detection and why we're disabling Turnitin's AI detector (16 August 2023) : https://www.vanderbilt.edu/brightspace/2023/08/16/guidance-on-ai-detection-and-why-were-disabling-turnitins-ai-detector/

University of Sydney, False flags and broken trust: can we tell if AI has been used? (30 October 2025) : https://educational-innovation.sydney.edu.au/teaching@sydney/false-flags-and-broken-trust-can-we-tell-if-ai-has-been-used/

Frequently Asked Questions

Are AI writing detectors biased against non-native English speakers?

Yes, and the effect has been measured. Liang and colleagues found that seven detectors falsely flagged 61.22 percent of TOEFL essays written by humans, while classifying essays by American eighth-graders almost perfectly. The mechanism is perplexity scoring, which treats predictable word choice as machine-like, and second-language writers use a narrower, more predictable vocabulary by design. Newer detectors report better numbers on ESL corpora, but the underlying signal being measured has not changed.

Can a watermark prove who wrote a document?

No. Anthropic's documentation for Claude's SynthID-Text watermark, published on 14 August 2026, states that the watermark can only show the model was likely involved at some point, and that it cannot distinguish text Claude wrote from text Claude heavily edited. A negative result is weaker still. Unwatermarked text may be human, or from a model that does not watermark, or watermarked output that was rewritten or translated past the signal.

What should a university use instead of an AI detector?

Process evidence rather than text analysis. Timestamped version history, a short oral defence in which the student explains decisions inside their own paper, drafts and supervision notes, and a disclosure policy carrying no penalty so that people actually use it. Turnitin's own guidance says the score exists to start a conversation, not to conclude one. The honest objection is cost: oral defence at scale needs staff hours that automated scoring does not.