All insights
Instrument plate showing 79% mounted on a dimension line, with 96% marked above it and the interval between them shaded
The TechnologyApril 1, 20267 min read

The Detector Is a Tax on Arab Students

AI-detection tools read human-written English correctly 96% of the time. On human writing machine-translated into English, they fall to 79%. Arab students who draft in Arabic pay a penalty for the language they think in.

An AI detector reads a genuinely human-written English essay correctly about 96% of the time. Take a genuinely human-written essay in another language, machine-translate it into English, and detector accuracy falls to 79%.

Seventeen points, for the act of translating.

That is not a rounding error in a marginal tool. It is the average across fourteen detection systems, measured in a peer-reviewed study, and it describes exactly what a Jordanian, Egyptian or Saudi undergraduate does when they draft an assignment in Arabic and submit it in English. The detector is not reading for dishonesty. It is reading for fluency, and it is charging a penalty to the students who have to translate to be understood.

The tools do not work, and their own numbers say so

Weber-Wulff and colleagues tested fourteen systems: twelve publicly available tools plus two commercial products, Turnitin and PlagiarismCheck. They published the results in the International Journal for Educational Integrity in 2023. Their verdict, verbatim, is that the tools "are neither accurate nor reliable (all scored below 80% of accuracy and only 5 over 70%)."

The accuracy breakdown by document class is the part worth memorising:

Document type Average accuracy
Human-written 96%
Human-written, machine-translated into English 79%
AI-generated, unmodified 75–77%
AI-generated, then edited by a human 42–45%
AI-generated, then machine-paraphrased 26–31%

Read the table from the bottom up and the product is worse than useless. A student who runs generated text through a paraphraser is caught roughly a quarter to a third of the time. A student who wrote the essay themselves in Arabic and translated it is wrongly caught far more often than one who wrote it in English. The best tool in the set was Turnitin at 81%. The worst, Content at Scale, scored 43%, which is worse than a coin flip on a two-class problem.

The authors also found that every tool but one was more accurate at recognising human writing than AI writing. The classifiers are tuned, they suggest probably deliberately, to err toward calling text human. That tuning is what keeps the headline false-positive rate low. It does not survive contact with writing that does not sound native.

The second failure stacks on the first

Liang and colleagues at Stanford ran seven widely used GPT detectors over 91 TOEFL essays written by humans and 88 US eighth-grade essays written by humans, and published the result in Patterns in July 2023. Every essay in both sets was human-written. There was nothing to detect.

On the American eighth-graders, the detectors were close to perfect: an average false positive rate of 5.19%. On the TOEFL essays, the average false positive rate was 61.22%. Eighteen of the 91 TOEFL essays, 19.78%, were flagged as machine-written by all seven detectors at once. Eighty-nine of the 91, or 97.80%, were flagged by at least one.

The mechanism is perplexity. Essays that all seven detectors agreed were AI had significantly lower perplexity than the rest, at p = 9.74E-05. Low perplexity means predictable word choice, and predictable word choice is what constrained second-language writing looks like. The detector is measuring vocabulary range and calling it fraud.

Then the authors ran the experiment that ends the argument. They asked a language model to enrich the vocabulary of the same TOEFL essays. The false positive rate fell from 61.22% to 11.77%, and only one of the 91 essays remained unanimously flagged. Running the manipulation in reverse — asking a model to simplify the American essays "as if written by a non-native speaker" — pushed misclassification of those native-written essays from 5.19% up to 56.65%.

A detector that a single prompt defeats does not deter anyone. It only sorts by fluency.

Both studies are small. Ninety-one essays and 88 essays, seven tools in one, fourteen in the other, all measured in 2023 against models that have since been replaced. That is a real limitation and it should be stated. It cuts one way only: nobody has since published a study showing the bias has gone.

The vendors and the universities already acted

OpenAI shut down its own detector on 20 July 2023, citing its low rate of accuracy. The published performance was 26% true positives against a 9% false-positive rate on human text. The company with the most direct commercial interest in a working detector built one, measured it, and withdrew it.

Vanderbilt University disabled Turnitin's AI detector on 16 August 2023. Its reasoning was arithmetic. Turnitin claimed a 1% false positive rate; Vanderbilt had submitted 75,000 papers in 2022; one percent of that is roughly 750 student papers wrongly accused in a single year at a single institution. Vanderbilt also cited the non-native-speaker research directly, along with Turnitin's refusal to explain how the detector worked.

Neither decision is new. Both are three years old.

Why this lands hardest here

The modal Arab university student writing an English-language assignment is two things at once: a non-native writer, and frequently a translator. The Stanford finding applies to the first. The Weber-Wulff finding applies to the second. Nobody has measured what happens when they stack, because nobody has run either study on Arabic-first writers.

That missing study is the point. There is no published false-positive rate for detectors applied to Arab students' English writing. Every university in the Gulf and the Levant currently running detection is operating on evidence collected from Chinese, Spanish and other TOEFL cohorts, and extrapolating.

The one piece of regional evidence points the same way. Researchers at Prince Mohammad Bin Fahd University in Saudi Arabia submitted handwritten examination papers — physically impossible to have been produced by a chatbot — to AI-detection programs, and found an association between higher exam grades and greater false detection. The better a student wrote, the more likely the software called them a cheat. It is a small, single-institution conference paper with limited peer review, and on its own it proves nothing. It is also one of very few Gulf-based data points that exist, and it agrees with the two large studies.

There is a regulatory dimension that most vendors selling into the region have not priced. The EU AI Act names systems "intended to be used for monitoring and detecting prohibited behaviour of students during tests" as high-risk under Annex III. Those obligations were pushed from August 2026 to 2 December 2027 under the Digital Omnibus agreement, but they were deferred, not softened. Any detection product with European customers, or with a Gulf customer that mirrors European requirements in procurement, inherits documentation, accuracy and human-oversight duties that the current generation of tools cannot evidence.

What to do about it

If you run a university in the region, turn the detector off. Vanderbilt published the reasoning in 2023 and nothing since has undermined it. Running detection on a student body that drafts in Arabic means accepting a false-accusation rate you have never measured, on a population you have never tested, using a tool whose own vendor cannot describe its method. An academic-integrity case built on a detector score is a case built on a number nobody can defend under cross-examination.

If you are building assessment software, do not build detection. The category is a trap: the false-positive rate rises with the honesty of the student and falls with the sophistication of the cheat. Build for process evidence instead — drafting history, oral defence, supervised writing, work that is produced where it can be observed. Those methods do not degrade when the model improves.

If you sell a detector into MENA, publish per-language false-positive rates. Nobody has. The first vendor to release false-positive rates for Arabic-first English writers, with the test set open, would be making a genuine contribution and would also be the only one able to answer the question a procurement officer should be asking.

If you are pitching us and detection is in the deck, expect this conversation. [FIKR TO CONFIRM: whether the fund wants to state a formal position that it will not back AI-detection products, or keep this as a diligence question.]

The honest summary is short. Detection tools flag honest students who translate, and miss dishonest students who paraphrase. The bias against non-native writers disappears with one prompt, which means the tool disciplines the students who did not think to ask. In a region where hundreds of thousands of students sit English-language assessments in their second language, that is not a technical shortcoming. It is a transfer of risk from the institution to the student, and the students paying it are the ones who did the work.

Sources

  1. 0114 detection systems tested (12 public tools plus Turnitin and PlagiarismCheck); all scored below 80% accuracy and only 5 above 70%; accuracy 96% on human-written text, 79% on human-written text machine-translated into English, 75–77% on unmodified AI text, 42–45% on AI text with human editing, 26–31% on AI text with machine paraphrase; Turnitin best at 81%, Content at Scale 43%; every tool except GPT Zero was more accurate on human text than on AI text — Weber-Wulff et al., International Journal for Educational Integrity 19:26 (Springer Nature), June 1, 2023
  2. 027 GPT detectors evaluated on 91 human-authored TOEFL essays and 88 human-authored US 8th-grade essays; false positive rate 61.22% on TOEFL essays against 5.19% on US 8th-grade essays; 18 of 91 (19.78%) flagged unanimously by all seven; 89 of 91 (97.80%) flagged by at least one; unanimously flagged essays had significantly lower perplexity (p = 9.74E-05) — Liang, Yuksekgonul, Mao, Wu & Zou, Patterns (Cell Press), Stanford University, July 1, 2023
  3. 03Prompting a model to enrich the TOEFL essays' language dropped the false positive rate from 61.22% to 11.77%, leaving 1 of 91 (1.10%) unanimously flagged; prompting it to simplify word choices as a non-native speaker would raised misclassification of native 8th-grade essays from 5.19% to 56.65% — Liang, Yuksekgonul, Mao, Wu & Zou, Patterns (Cell Press), July 1, 2023
  4. 04OpenAI withdrew its AI Text Classifier on 20 July 2023 'due to its low rate of accuracy'; the tool correctly identified 26% of AI-written text and falsely flagged human-written text 9% of the time — OpenAI, July 20, 2023
  5. 05Vanderbilt University disabled Turnitin's AI detector on 16 August 2023; against Turnitin's claimed 1% false positive rate and the 75,000 papers Vanderbilt submitted in 2022, roughly 750 student papers would have been incorrectly flagged — Vanderbilt University, Office of Learning Innovation, August 16, 2023
  6. 06Handwritten examinations from an introductory academic-writing class at Prince Mohammad Bin Fahd University, Saudi Arabia, submitted to AI-detection programs, showed an association between higher exam grades and greater false detection of AI-generated writing — Waked, Abdelsalem, Ashraf, Pilotti & El Alaoui, ICRES 2024 Proceedings, pp. 573–586, January 1, 2024
  7. 07EU AI Act Annex III point 3(d) classifies AI systems intended for monitoring and detecting prohibited student behaviour during tests as high-risk — EU AI Act (Annex III), entered into force 1 August 2024, August 1, 2024
  8. 08Under the Digital Omnibus political agreement, stand-alone Annex III high-risk obligations — explicitly including education — were postponed from 2 August 2026 to 2 December 2027 — Gibson Dunn, May 27, 2026