An AI detector reads a genuinely human-written English essay correctly about 96% of the time. Take a genuinely human-written essay in another language, machine-translate it into English, and detector accuracy falls to 79%.
Seventeen points, for the act of translating.
That is not a rounding error in a marginal tool. It is the average across fourteen detection systems, measured in a peer-reviewed study, and it describes exactly what a Jordanian, Egyptian or Saudi undergraduate does when they draft an assignment in Arabic and submit it in English. The detector is not reading for dishonesty. It is reading for fluency, and it is charging a penalty to the students who have to translate to be understood.
The tools do not work, and their own numbers say so
Weber-Wulff and colleagues tested fourteen systems: twelve publicly available tools plus two commercial products, Turnitin and PlagiarismCheck. They published the results in the International Journal for Educational Integrity in 2023. Their verdict, verbatim, is that the tools "are neither accurate nor reliable (all scored below 80% of accuracy and only 5 over 70%)."
The accuracy breakdown by document class is the part worth memorising:
| Document type | Average accuracy |
|---|---|
| Human-written | 96% |
| Human-written, machine-translated into English | 79% |
| AI-generated, unmodified | 75–77% |
| AI-generated, then edited by a human | 42–45% |
| AI-generated, then machine-paraphrased | 26–31% |
Read the table from the bottom up and the product is worse than useless. A student who runs generated text through a paraphraser is caught roughly a quarter to a third of the time. A student who wrote the essay themselves in Arabic and translated it is wrongly caught far more often than one who wrote it in English. The best tool in the set was Turnitin at 81%. The worst, Content at Scale, scored 43%, which is worse than a coin flip on a two-class problem.
The authors also found that every tool but one was more accurate at recognising human writing than AI writing. The classifiers are tuned, they suggest probably deliberately, to err toward calling text human. That tuning is what keeps the headline false-positive rate low. It does not survive contact with writing that does not sound native.
The second failure stacks on the first
Liang and colleagues at Stanford ran seven widely used GPT detectors over 91 TOEFL essays written by humans and 88 US eighth-grade essays written by humans, and published the result in Patterns in July 2023. Every essay in both sets was human-written. There was nothing to detect.
On the American eighth-graders, the detectors were close to perfect: an average false positive rate of 5.19%. On the TOEFL essays, the average false positive rate was 61.22%. Eighteen of the 91 TOEFL essays, 19.78%, were flagged as machine-written by all seven detectors at once. Eighty-nine of the 91, or 97.80%, were flagged by at least one.
The mechanism is perplexity. Essays that all seven detectors agreed were AI had significantly lower perplexity than the rest, at p = 9.74E-05. Low perplexity means predictable word choice, and predictable word choice is what constrained second-language writing looks like. The detector is measuring vocabulary range and calling it fraud.
Then the authors ran the experiment that ends the argument. They asked a language model to enrich the vocabulary of the same TOEFL essays. The false positive rate fell from 61.22% to 11.77%, and only one of the 91 essays remained unanimously flagged. Running the manipulation in reverse — asking a model to simplify the American essays "as if written by a non-native speaker" — pushed misclassification of those native-written essays from 5.19% up to 56.65%.
A detector that a single prompt defeats does not deter anyone. It only sorts by fluency.
Both studies are small. Ninety-one essays and 88 essays, seven tools in one, fourteen in the other, all measured in 2023 against models that have since been replaced. That is a real limitation and it should be stated. It cuts one way only: nobody has since published a study showing the bias has gone.
The vendors and the universities already acted
OpenAI shut down its own detector on 20 July 2023, citing its low rate of accuracy. The published performance was 26% true positives against a 9% false-positive rate on human text. The company with the most direct commercial interest in a working detector built one, measured it, and withdrew it.
Vanderbilt University disabled Turnitin's AI detector on 16 August 2023. Its reasoning was arithmetic. Turnitin claimed a 1% false positive rate; Vanderbilt had submitted 75,000 papers in 2022; one percent of that is roughly 750 student papers wrongly accused in a single year at a single institution. Vanderbilt also cited the non-native-speaker research directly, along with Turnitin's refusal to explain how the detector worked.
Neither decision is new. Both are three years old.
Why this lands hardest here
The modal Arab university student writing an English-language assignment is two things at once: a non-native writer, and frequently a translator. The Stanford finding applies to the first. The Weber-Wulff finding applies to the second. Nobody has measured what happens when they stack, because nobody has run either study on Arabic-first writers.
That missing study is the point. There is no published false-positive rate for detectors applied to Arab students' English writing. Every university in the Gulf and the Levant currently running detection is operating on evidence collected from Chinese, Spanish and other TOEFL cohorts, and extrapolating.
The one piece of regional evidence points the same way. Researchers at Prince Mohammad Bin Fahd University in Saudi Arabia submitted handwritten examination papers — physically impossible to have been produced by a chatbot — to AI-detection programs, and found an association between higher exam grades and greater false detection. The better a student wrote, the more likely the software called them a cheat. It is a small, single-institution conference paper with limited peer review, and on its own it proves nothing. It is also one of very few Gulf-based data points that exist, and it agrees with the two large studies.
There is a regulatory dimension that most vendors selling into the region have not priced. The EU AI Act names systems "intended to be used for monitoring and detecting prohibited behaviour of students during tests" as high-risk under Annex III. Those obligations were pushed from August 2026 to 2 December 2027 under the Digital Omnibus agreement, but they were deferred, not softened. Any detection product with European customers, or with a Gulf customer that mirrors European requirements in procurement, inherits documentation, accuracy and human-oversight duties that the current generation of tools cannot evidence.
What to do about it
If you run a university in the region, turn the detector off. Vanderbilt published the reasoning in 2023 and nothing since has undermined it. Running detection on a student body that drafts in Arabic means accepting a false-accusation rate you have never measured, on a population you have never tested, using a tool whose own vendor cannot describe its method. An academic-integrity case built on a detector score is a case built on a number nobody can defend under cross-examination.
If you are building assessment software, do not build detection. The category is a trap: the false-positive rate rises with the honesty of the student and falls with the sophistication of the cheat. Build for process evidence instead — drafting history, oral defence, supervised writing, work that is produced where it can be observed. Those methods do not degrade when the model improves.
If you sell a detector into MENA, publish per-language false-positive rates. Nobody has. The first vendor to release false-positive rates for Arabic-first English writers, with the test set open, would be making a genuine contribution and would also be the only one able to answer the question a procurement officer should be asking.
If you are pitching us and detection is in the deck, expect this conversation. [FIKR TO CONFIRM: whether the fund wants to state a formal position that it will not back AI-detection products, or keep this as a diligence question.]
The honest summary is short. Detection tools flag honest students who translate, and miss dishonest students who paraphrase. The bias against non-native writers disappears with one prompt, which means the tool disciplines the students who did not think to ask. In a region where hundreds of thousands of students sit English-language assessments in their second language, that is not a technical shortcoming. It is a transfer of risk from the institution to the student, and the students paying it are the ones who did the work.




