AI Detection in Education Cannot Work as Policy
Francis Okafor
On this page
- Assessment was always measuring a proxy
- Why AI detection in education fails as policy
- Assessment that does not trust the artifact
- Two students, one rubric, different exams
- The specific risk to countries building an engineering workforce
- The strongest case against everything I just said
- What the detector was actually measuring
- Sources
I have had the same message more than once. A student I mentor, writing at some late hour, sends a screenshot of an integrity flag on work she produced herself. English is her second language. Sometimes her third. She wants to know what evidence she can put in front of a panel to prove she wrote her own sentences, and the honest answer is that there is none, because you cannot prove authorship of prose after the fact.
AI detection in education has been argued about for close to four years now, almost entirely as a cheating story. Who cheated. How many. What percentage. It is the least interesting version of the argument and it has taken nearly all the oxygen.
Here is the version I think is load-bearing. Assessment has never measured the thing it cares about. It measures a proxy. The essay, the problem set, the report. We grade the artifact and infer the capability, and that inference held for a few centuries because producing a convincing artifact without the capability was expensive. It stopped being expensive. The proxy broke. The capability did not move a millimetre.
Everything else follows from that.
Assessment was always measuring a proxy
Nobody wants the essay. What the marker wants to know is whether this person can hold a position, anticipate an objection and change their mind in public without falling apart. The essay was a container for that. Cheap to collect, portable, gradeable at scale.
Same with the problem set. Nobody needs another solved integral. The question was whether the student can carry a derivation from one end to the other without losing the thread.
The proxy always leaked. Older siblings, past papers, ghostwriters, essay mills that charged by the word and delivered in three days. Institutions tolerated the leakage because it carried a price and a delay, and price and delay are a filter. Generative tools removed both. That is the entire change. Not the arrival of dishonesty, which is old, but the collapse of the cost floor underneath it.
So a mark on a submitted artifact now carries almost no information about the student. It carries information about the artifact. Those used to be nearly the same claim. They have come apart, and detection cannot put them back together, because detection tries to repair the inference from the artifact side and the artifact side is exactly where the damage is.
Simplify the word choices in a genuine American schoolchild's essay and detector misclassification jumps from 5.19 percent to 56.65 percent. Same author, same work, style edited.
Why AI detection in education fails as policy
The clearest evidence is a 2023 study by Weixin Liang and colleagues, published in Patterns and posted as arXiv:2304.02819. Seven widely used GPT detectors were run over 91 TOEFL essays written by non-native English speakers, sourced from a Chinese educational forum, and 88 essays by US eighth-graders taken from the Hewlett Foundation's Automated Student Assessment Prize dataset.
The detectors were near-perfect on the American children. On the TOEFL essays the average false positive rate was 61.22 percent. At least one detector flagged 89 of the 91. All seven agreed on 18 of them, so roughly one student in five would have been convicted by unanimous machine verdict.
Then the finding that should have ended the argument. The researchers had ChatGPT rewrite the same essays with the instruction to enhance the word choices to sound more like a native speaker. The false positive rate fell to 11.77 percent. Run the experiment backwards and it is worse: when they simplified the word choices in the genuine American essays to mimic a non-native writer, misclassification rose from 5.19 percent to 56.65 percent. Same authors, same work, style edited. The detector flipped.
These tools score perplexity, roughly how surprising each next word is. Low surprise reads as machine. It also reads as a careful second-language writer with a smaller active vocabulary and a preference for constructions she knows are correct. I write English that way. Most of the world does.
Vanderbilt University disabled Turnitin's AI detector in August 2023 and published its reasoning. The post takes Turnitin's claimed one percent false positive rate, applies it to roughly 75,000 annual submissions and arrives at about 750 papers wrongly labelled. It also states that Turnitin gives no detailed information about how the determination is made. That one percent is a vendor figure with no published methodology behind it, and it has been quoted in institutional policy ever since as though it were a measured constant. It is not one.
Common Sense Media's 2024 report, The Dawn of the AI Era, surveyed by Ipsos across 1,045 US teenagers, found one in ten saying a teacher had flagged their work as AI-generated when it was not. Twenty percent of Black teens, against 10 percent of Latino and 7 percent of white teens. Those are self-reports rather than measured detector error and the distinction matters. For the student sitting in the meeting, the operative question is who gets accused.
Assessment that does not trust the artifact
The University of Sydney published its answer in October 2025. Adam Bridgeman and Danny Liu describe a two-lane model. Lane one is secured assessment, defined as in-person supervised work used to validate learning, and it includes interactive oral conversations, placements and in-person practical work rather than only written examinations. Lane two is open assessment, which carries the majority of the load, where contemporary tools are assumed and scaffolded into the unit. Their justification is blunt. They estimate as much as 85 percent of their current assessment could be at least satisfactorily completed by a moderately tech-savvy novice.
For engineering the translation is easy, because industry solved this before education had to. I do not need to know who typed the code. I need to know whether you can say why you chose that structure, what fails at ten times the load and what you would delete first. Fifteen minutes at a whiteboard settles it and no detector is involved.
The instruments are unglamorous. Short oral defence of submitted work, low ceremony, on a sample rather than everyone. Supervised in-class production for the components that must be secured. Version history, commit messages, draft trails that show a position changing. Not because these cannot be gamed, they can, but because the cheapest way to pass an oral defence is to understand the material, which is the outcome the whole apparatus exists to produce.
Every hiring process I have watched over the last two years has quietly shifted weight off the take-home exercise and onto the live conversation. Education is running behind the labour market it feeds.
Two students, one rubric, different exams
The Higher Education Policy Institute's Student Generative AI Survey, published in February 2025, found 88 percent of UK students had used generative AI for assessments, up from 53 percent a year earlier. Overall use went from 66 to 92 percent. Only 36 percent said their institution had given them any AI skills training. HEPI also recorded a widening divide, with more socio-economically advantaged students, male students and those in STEM and health subjects more likely to be using these tools.
That last finding should worry examination boards more than the cheating numbers do, and it gets a fraction of the attention.
Consider the arithmetic. OpenAI's consumer ladder now runs from a free tier through Go at eight dollars a month, Plus at twenty and Pro at one hundred or two hundred. Wikipedia's country comparison of minimum wages puts Nigeria's national minimum at seventy thousand naira a month, worth roughly forty-four US dollars after the naira's devaluation. Plus therefore costs about 45 percent of a full month at Nigeria's minimum wage, and the cheaper Go tier, sold in Nigeria at seven thousand naira, still takes a tenth of it. The upper tier costs four and a half months.
Two students submit against the same rubric. One has a frontier model, long context and no rate limits. The other queues on a free tier, on a phone, on metered data, at whatever hour the servers are quietest. They are marked identically and the transcript records none of it. There has always been an equipment gap in education: a library card, a quiet room, a parent who reads. The difference now is that the gap sits inside the act of production and it is priced monthly.
The specific risk to countries building an engineering workforce
I care about this from a particular position. I moved to China in 2018 and have spent eight years inside its technology ecosystem, and a good deal of my time goes to the connective work between that ecosystem and African engineering talent. Two failure modes are live at once and they pull in opposite directions.
The first is over-policing. Most of anglophone Africa teaches and examines in English, which is a second or third language for most of the students sitting the papers. Import a detector calibrated on native prose into that setting and Liang's 61.22 percent lands squarely on the population the institution exists to certify. A university that adopts detection as policy will manufacture a false accusation rate concentrated on its own students, read that rate as evidence of an epidemic and tighten the policy further.
The second is under-teaching, and it is slower and worse. An engineering credential has no intrinsic value. It is a claim that somebody competent checked. If the checking was done on artifacts that no longer carry information, the claim is empty, and employers work that out within a hiring cycle or two. Then they build their own screens. Screens are expensive, and expensive screens get run on graduates from institutions the employer already trusts. That list is short.
A country trying to scale its engineering workforce by an order of magnitude cannot afford either outcome. It cannot afford to falsely accuse the students it needs, and it cannot afford to issue degrees the market has stopped reading.
The strongest case against everything I just said
Two objections, both serious, and I do not think either has been answered.
The first is that some knowledge genuinely has to sit in a person's head. Not retrievable, resident. A structural engineer who cannot feel that a load figure is wrong will never think to check it. A clinician mid-procedure has no lookup window. Reasoning happens in working memory and working memory has no room for a search query, so material you have to look up is material you cannot think with. Closed-book examination is not nostalgia. It is the only instrument that tests what a person carries unaided, and the process-evidence camp tends to skate past this because its advocates mostly assess writing rather than judgement under load.
The second objection is arithmetic and it is fatal to the naive version of my position. Four hundred students. Two staff. Fifteen minutes of oral defence each is one hundred hours of examination per round, before marking, before moderation, before appeals. Draft trails are worse, because reading a version history takes longer than reading a final. Much of the assessment reform literature was written by people with seminar-sized cohorts and light marking loads, and it does not survive contact with a first-year service course in an underfunded university.
The partial answer, and I want to be clear that it is partial, is to stop trying to secure everything. Sydney's model secures a minority of assessment and assures outcomes across a pathway rather than at every point. Sample the orals. Secure two or three checkpoints per degree, not per unit. That cuts the bill by an order of magnitude without eliminating it. Somebody still has to pay, and in the systems most exposed to all of this the money is not there.
What the detector was actually measuring
Go back to the 61.22 percent, and to what happened when it fell to 11.77, and to the American essays that jumped from 5.19 to 56.65 once their vocabulary was simplified.
The detectors were never measuring honesty. They were measuring distance from an average of native-speaker prose. A student who writes plainly, in constructions she has checked are correct, with a vocabulary earned late rather than absorbed early, registers as synthetic. A well-resourced student with a good prompt registers as human. The instrument inverts on precisely the axis you would least want it to.
We built a test that a funded cheat passes and an honest second-language writer fails, then printed the output in the student handbook under the heading of integrity.
The alternative is watching people work, which is the oldest form of assessment there is and also the most expensive. The institutions most exposed to the false accusation problem are, with grim reliability, the ones least able to fund the fix. So the distance keeps opening between what an institution can afford to measure and what it actually needs to know. What sits in that gap is trust, which is the one thing this argument has already spent.
Sources
Liang, Yuksekgonul, Mao, Wu and Zou, GPT detectors are biased against non-native English writers, Patterns (arXiv:2304.02819): https://arxiv.org/abs/2304.02819
Vanderbilt University, Guidance on AI detection and why we're disabling Turnitin's AI detector, 16 August 2023: https://www.vanderbilt.edu/brightspace/2023/08/16/guidance-on-ai-detection-and-why-were-disabling-turnitins-ai-detector/
Common Sense Media, The Dawn of the AI Era: Teens, Parents, and the Adoption of Generative AI at Home and School, 2024: https://www.commonsensemedia.org/research/the-dawn-of-the-ai-era-teens-parents-and-the-adoption-of-generative-ai-at-home-and-school
HEPI, Student Generative AI Survey 2025, February 2025: https://www.hepi.ac.uk/2025/02/26/student-generative-ai-survey-2025/
Bridgeman and Liu, Two parallel lanes: the roadmap for a future-ready transformative education, University of Sydney, October 2025: https://educational-innovation.sydney.edu.au/teaching@sydney/two-parallel-lanes-the-roadmap-for-a-future-ready-transformative-education/
Center for Democracy and Technology, Late applications: disproportionate effects of generative AI detectors on English learners, December 2023: https://cdt.org/insights/brief-late-applications-disproportionate-effects-of-generative-ai-detectors-on-english-learners/
Wikipedia, ChatGPT (subscription tiers and pricing): https://en.wikipedia.org/wiki/ChatGPT
Wikipedia, List of countries by minimum wage (Nigeria): https://en.wikipedia.org/wiki/List_of_countries_by_minimum_wage
Frequently Asked Questions
Are AI detectors accurate enough to use as evidence in an academic integrity case?
No. In the 2023 Patterns study by Liang and colleagues (arXiv:2304.02819), seven widely used detectors produced an average 61.22 percent false positive rate on 91 TOEFL essays written by humans, and all seven unanimously flagged 18 of them. Vanderbilt University disabled Turnitin's AI detector in August 2023 partly because the vendor publishes no methodology behind its claimed one percent false positive rate. A score with no published error model and no per-document confidence interval is not evidence.
Why do AI detectors flag non-native English writers so often?
They score perplexity, a measure of how surprising each next word is. Non-native writers tend to use a smaller active vocabulary and safer constructions, which lowers perplexity and looks machine-generated. Liang and colleagues demonstrated the mechanism directly: asking ChatGPT to make TOEFL essays sound more native dropped the false positive rate from 61.22 to 11.77 percent, and simplifying genuine US eighth-grade essays to mimic a non-native writer pushed misclassification from 5.19 to 56.65 percent.
What replaces the essay if the submitted artifact cannot be trusted?
Direct observation of the capability rather than the product. The University of Sydney's two-lane model, published by Adam Bridgeman and Danny Liu in October 2025, secures a minority of assessment through in-person supervised work including interactive oral conversations, placements and practical work, while treating the majority as open assessment where tool use is assumed and taught. Short oral defence on a sample, supervised in-class production and version history are the usual instruments.
Does banning AI tools fix the equity problem?
No, it usually deepens it. HEPI's February 2025 survey found 88 percent of UK students using generative AI for assessments, up from 53 percent, with use skewed towards more socio-economically advantaged students. A ban shifts the contest to who can conceal use, which correlates with who can afford better tools. It also concentrates enforcement on the students detectors misread, which are disproportionately non-native English writers.
How do you run oral assessment with four hundred students and two staff?
You do not run it on every task. Fifteen minutes each is one hundred hours per round, which no department has. The workable version secures a small number of checkpoints at programme level rather than per unit, samples orals instead of examining everyone and accepts that some assessment stays open. It still costs money that many systems do not have, and any reform proposal that skips that line is not a proposal.