Four researchers presenting at last year's Educational Data Mining conference ran a straightforward experiment and got an answer nobody in this sector wants. They took 252 learners, gave them twelve mathematics problems, and split them into four groups. One group got feedback that was always correct. One got feedback that was wrong half the time. One got feedback that was wrong every single time. The last group got no feedback at all.
Then they asked the learners how useful the feedback had been.
The group that received feedback that was wrong on every problem rated it as more useful than the group that received feedback that was always right.
The study was preregistered, which means the authors filed their analysis plan before they collected any data. That matters here: it rules out the possibility that this result was found by searching the data for something surprising after the fact. It was the question they set out to answer, and this is the answer they got.
What the trial found, in full
The headline result is the one above, but the rest of the paper is not what you would guess from it.
Learners in the all-wrong condition did learn. Their scores improved significantly against the group that got no feedback at all — as did the scores of the group whose feedback was always correct. Getting a response, even a wrong one, beat getting nothing. That is a finding about attention and effort, not about accuracy, and it is worth stating plainly rather than burying: the all-wrong group also spent more time on task and reported more confusion than the control.
Learners were not entirely fooled, either. When asked how accurate the feedback was, their ratings dropped as the error rate climbed. They could feel that something was off.
They just did not act on it. Usefulness ratings did not follow accuracy ratings down. And the students who reliably identified the erroneous feedback as inaccurate were the ones who already knew the material well. Everyone else took the wrong answer and rated it highly.
Read that last part slowly, because it is the commercial finding. The students least able to detect an error are exactly the low-prior-knowledge students the product is sold to help. A tutoring product that is wrong will look best to the students who need it most.
Why satisfaction is the wrong instrument
Almost every EdTech company measures student satisfaction. It goes in the deck, in the data room, and in the renewal conversation with a school. "Students love it" is the most common piece of evidence offered for an AI tutor, and after this trial it is worth precisely nothing as evidence of correctness.
The in-product usage numbers are no better, and there is a cleaner experiment on that.
Bastani and colleagues ran a field trial with nearly 1,000 high school mathematics students in Turkey, in three groups. One got a standard chatbot. One got a version with prompts written to protect learning — hints instead of answers, in effect. One got a textbook and their notes.
During practice, with the tools in front of them, the chatbot group performed 48% better than the textbook group and the safeguarded-tutor group performed 127% better. Those are enormous numbers, and they are the numbers an EdTech dashboard would report.
Then the students sat an exam with the AI taken away. The chatbot group scored 17% worse than the students who had never had access at all. The safeguarded group scored about the same as the control — no harm, and no gain either.
So the 127% improvement was almost entirely an artefact of having the tool in the room. Any metric measured inside the product was measuring the tool's performance, not the student's.
There is now a third line of evidence pointing the same way, from a different direction entirely. A team at UC Irvine, working with McGraw Hill, examined a decade of activity on the ALEKS mathematics platform: 3.2 million learning interactions and 12.2 million placement-assessment response times. They compared problem types a chatbot can help with, such as text word problems, against problem types it cannot, such as interactive graph problems. After ChatGPT's release, time spent on the AI-susceptible problems fell 2.8% per quarter among college students, compounding to −26.9% over eleven quarters. High schoolers fell 31.3%. Middle schoolers 9.0%. Fifth graders, no detectable change. On proctored retention items, the odds of answering correctly fell 25% — and the whole divergence disappears under proctoring, which rules out the comfortable explanation that students had simply become more efficient. This one is a preprint and it is industry-authored, so treat it as strong design rather than settled fact. But it is the third independent measurement in which faster, happier, more engaged use of an AI tool came with less learning.
Why your tutor will be wrong on exactly the problems teachers write
The natural response to all of this is that the errors are a temporary problem. Better models, fewer hallucinations, and the ratings become meaningful again.
There is a specific reason to doubt that, and it comes from a paper by Apple's machine-learning group.
The standard way to test whether a model can do school mathematics is a benchmark of grade-school word problems. The Apple team rebuilt that benchmark as a set of templates, so the same problem could be reissued with different names and different numbers — the same reasoning, new surface. Three things happened. Model scores varied noticeably across versions of what was mathematically the same question. Scores fell when only the numbers were changed. Scores fell further as more clauses were added to the problem.
Then they added a single clause that looked relevant but contributed nothing to the answer. Performance dropped by up to 65% across every current model they tested.
Now think about what a real word problem written by a real teacher looks like. It has names, a context, a distracting detail, a number that does not go anywhere. Irrelevant-but-plausible clauses are not an adversarial edge case in education. They are the ordinary furniture of a curriculum question. The failure mode the Apple team induced deliberately is the failure mode a tutoring product will meet on a Tuesday.
Grading has a related problem. A March 2026 preprint compared GPT- and Llama-family models against human graders on student essays and found agreement "remains relatively weak," with a consistent direction to the error: models over-score short, undeveloped essays and under-score longer essays that contain small spelling or grammar mistakes. Worse, the models' scores are internally consistent with the feedback they themselves generate. A grader that is wrong in a stable, self-justifying way is harder to catch than one that is noisy.
What this means in this region
None of the studies above were run in an Arab classroom. That is its own finding, and we have written about it elsewhere.
What is documented regionally is the shape of the exposure. EdTech Hub's January 2026 policy brief on AI in education across MENA found adoption "concentrated among urban elites, private schools, and well-resourced universities," with public-sector schools less likely to adopt because of infrastructure gaps, higher student-teacher ratios and limited ministry-led integration. It also found that ministries in Jordan, Iraq and Egypt have not embedded AI-focused training into national teacher development.
Put the two together. The EDM trial says the students who cannot catch an error are the ones with the least prior knowledge. The EdTech Hub brief says the teachers who would catch it on their behalf have not been trained to. The regional pitch for AI tutoring is that it substitutes for teaching capacity that is not there — which is to say, it is aimed precisely at the classrooms with no second line of defence.
That is not an argument against building it. It is an argument for building it with the error rate measured, and for not treating a satisfaction score as a substitute for that measurement.
What a founder should do about it
Stop offering satisfaction scores as evidence of learning. They are evidence of experience, which is a real thing worth measuring for retention. They are not evidence that your product is right, and there is now a preregistered trial saying they run in the opposite direction.
Show the delta on an assessment the student takes without your product in the room. This is the single diligence question that separates the field, and most companies in it cannot answer. Bastani's exam is the template: same students, tool removed, compared against students who never had it.
Measure your error rate on your own curriculum, not on a benchmark. Take a hundred real questions from the syllabus you sell into, have a subject teacher mark your product's responses, and publish the rate. If the Apple result holds, your rate on messy teacher-written problems will be worse than your rate on clean benchmark ones, and you want to know that before a ministry does.
Treat the guardrails as the product. The safeguarded tutor in Turkey did no harm. The unguarded one made students measurably worse. The difference was prompt design and refusal behaviour, not model choice, and anyone can call the same model you are calling.
Segment your evidence by prior knowledge. If your gains come from students who already knew the material, you have built a revision tool and should sell it as one.
The ceiling here is worth remembering, because it is higher than this article makes it sound. A meta-analysis of 107 effect measurements across 14,321 participants found intelligent tutoring systems statistically indistinguishable from one-to-one human tutoring. Software can teach. The problem is not the category. The problem is that the metric the category has chosen to prove itself with rates a wrong answer above a right one.




