All insights
Typographic plate setting the article title in the Instrument Field system, with a usefulness rating scale drawn beneath it and the accurate and inaccurate conditions marked at the same point
The TechnologyMay 12, 20267 min read

Students Rate the Wrong Answers Higher

In a preregistered trial, 252 learners were given mathematics feedback that was wrong every single time. They rated it more useful than the feedback that was right, and only the students who already knew the material noticed.

Four researchers presenting at last year's Educational Data Mining conference ran a straightforward experiment and got an answer nobody in this sector wants. They took 252 learners, gave them twelve mathematics problems, and split them into four groups. One group got feedback that was always correct. One got feedback that was wrong half the time. One got feedback that was wrong every single time. The last group got no feedback at all.

Then they asked the learners how useful the feedback had been.

The group that received feedback that was wrong on every problem rated it as more useful than the group that received feedback that was always right.

The study was preregistered, which means the authors filed their analysis plan before they collected any data. That matters here: it rules out the possibility that this result was found by searching the data for something surprising after the fact. It was the question they set out to answer, and this is the answer they got.

What the trial found, in full

The headline result is the one above, but the rest of the paper is not what you would guess from it.

Learners in the all-wrong condition did learn. Their scores improved significantly against the group that got no feedback at all — as did the scores of the group whose feedback was always correct. Getting a response, even a wrong one, beat getting nothing. That is a finding about attention and effort, not about accuracy, and it is worth stating plainly rather than burying: the all-wrong group also spent more time on task and reported more confusion than the control.

Learners were not entirely fooled, either. When asked how accurate the feedback was, their ratings dropped as the error rate climbed. They could feel that something was off.

They just did not act on it. Usefulness ratings did not follow accuracy ratings down. And the students who reliably identified the erroneous feedback as inaccurate were the ones who already knew the material well. Everyone else took the wrong answer and rated it highly.

Read that last part slowly, because it is the commercial finding. The students least able to detect an error are exactly the low-prior-knowledge students the product is sold to help. A tutoring product that is wrong will look best to the students who need it most.

Why satisfaction is the wrong instrument

Almost every EdTech company measures student satisfaction. It goes in the deck, in the data room, and in the renewal conversation with a school. "Students love it" is the most common piece of evidence offered for an AI tutor, and after this trial it is worth precisely nothing as evidence of correctness.

The in-product usage numbers are no better, and there is a cleaner experiment on that.

Bastani and colleagues ran a field trial with nearly 1,000 high school mathematics students in Turkey, in three groups. One got a standard chatbot. One got a version with prompts written to protect learning — hints instead of answers, in effect. One got a textbook and their notes.

During practice, with the tools in front of them, the chatbot group performed 48% better than the textbook group and the safeguarded-tutor group performed 127% better. Those are enormous numbers, and they are the numbers an EdTech dashboard would report.

Then the students sat an exam with the AI taken away. The chatbot group scored 17% worse than the students who had never had access at all. The safeguarded group scored about the same as the control — no harm, and no gain either.

So the 127% improvement was almost entirely an artefact of having the tool in the room. Any metric measured inside the product was measuring the tool's performance, not the student's.

There is now a third line of evidence pointing the same way, from a different direction entirely. A team at UC Irvine, working with McGraw Hill, examined a decade of activity on the ALEKS mathematics platform: 3.2 million learning interactions and 12.2 million placement-assessment response times. They compared problem types a chatbot can help with, such as text word problems, against problem types it cannot, such as interactive graph problems. After ChatGPT's release, time spent on the AI-susceptible problems fell 2.8% per quarter among college students, compounding to −26.9% over eleven quarters. High schoolers fell 31.3%. Middle schoolers 9.0%. Fifth graders, no detectable change. On proctored retention items, the odds of answering correctly fell 25% — and the whole divergence disappears under proctoring, which rules out the comfortable explanation that students had simply become more efficient. This one is a preprint and it is industry-authored, so treat it as strong design rather than settled fact. But it is the third independent measurement in which faster, happier, more engaged use of an AI tool came with less learning.

Why your tutor will be wrong on exactly the problems teachers write

The natural response to all of this is that the errors are a temporary problem. Better models, fewer hallucinations, and the ratings become meaningful again.

There is a specific reason to doubt that, and it comes from a paper by Apple's machine-learning group.

The standard way to test whether a model can do school mathematics is a benchmark of grade-school word problems. The Apple team rebuilt that benchmark as a set of templates, so the same problem could be reissued with different names and different numbers — the same reasoning, new surface. Three things happened. Model scores varied noticeably across versions of what was mathematically the same question. Scores fell when only the numbers were changed. Scores fell further as more clauses were added to the problem.

Then they added a single clause that looked relevant but contributed nothing to the answer. Performance dropped by up to 65% across every current model they tested.

Now think about what a real word problem written by a real teacher looks like. It has names, a context, a distracting detail, a number that does not go anywhere. Irrelevant-but-plausible clauses are not an adversarial edge case in education. They are the ordinary furniture of a curriculum question. The failure mode the Apple team induced deliberately is the failure mode a tutoring product will meet on a Tuesday.

Grading has a related problem. A March 2026 preprint compared GPT- and Llama-family models against human graders on student essays and found agreement "remains relatively weak," with a consistent direction to the error: models over-score short, undeveloped essays and under-score longer essays that contain small spelling or grammar mistakes. Worse, the models' scores are internally consistent with the feedback they themselves generate. A grader that is wrong in a stable, self-justifying way is harder to catch than one that is noisy.

What this means in this region

None of the studies above were run in an Arab classroom. That is its own finding, and we have written about it elsewhere.

What is documented regionally is the shape of the exposure. EdTech Hub's January 2026 policy brief on AI in education across MENA found adoption "concentrated among urban elites, private schools, and well-resourced universities," with public-sector schools less likely to adopt because of infrastructure gaps, higher student-teacher ratios and limited ministry-led integration. It also found that ministries in Jordan, Iraq and Egypt have not embedded AI-focused training into national teacher development.

Put the two together. The EDM trial says the students who cannot catch an error are the ones with the least prior knowledge. The EdTech Hub brief says the teachers who would catch it on their behalf have not been trained to. The regional pitch for AI tutoring is that it substitutes for teaching capacity that is not there — which is to say, it is aimed precisely at the classrooms with no second line of defence.

That is not an argument against building it. It is an argument for building it with the error rate measured, and for not treating a satisfaction score as a substitute for that measurement.

What a founder should do about it

Stop offering satisfaction scores as evidence of learning. They are evidence of experience, which is a real thing worth measuring for retention. They are not evidence that your product is right, and there is now a preregistered trial saying they run in the opposite direction.

Show the delta on an assessment the student takes without your product in the room. This is the single diligence question that separates the field, and most companies in it cannot answer. Bastani's exam is the template: same students, tool removed, compared against students who never had it.

Measure your error rate on your own curriculum, not on a benchmark. Take a hundred real questions from the syllabus you sell into, have a subject teacher mark your product's responses, and publish the rate. If the Apple result holds, your rate on messy teacher-written problems will be worse than your rate on clean benchmark ones, and you want to know that before a ministry does.

Treat the guardrails as the product. The safeguarded tutor in Turkey did no harm. The unguarded one made students measurably worse. The difference was prompt design and refusal behaviour, not model choice, and anyone can call the same model you are calling.

Segment your evidence by prior knowledge. If your gains come from students who already knew the material, you have built a revision tool and should sell it as one.

The ceiling here is worth remembering, because it is higher than this article makes it sound. A meta-analysis of 107 effect measurements across 14,321 participants found intelligent tutoring systems statistically indistinguishable from one-to-one human tutoring. Software can teach. The problem is not the category. The problem is that the metric the category has chosen to prove itself with rates a wrong answer above a right one.

Sources

  1. 01Preregistered randomised controlled trial, N = 252 learners, twelve mathematics problem-solving tasks with a pre/post-test design and four conditions: 0% hallucinated feedback, 50%, 100%, and a no-feedback control. Significant learning gains in both the 0% and the 100% conditions relative to the no-feedback control; the 100% condition showed higher time on task and higher confusion than control; learners rated the 100%-hallucinated feedback as more useful than the accurate feedback; perceived accuracy declined as the hallucination rate rose but usefulness ratings did not track accuracy; only high-prior-knowledge learners reliably identified erroneous feedback as inaccurate — Steinbach, Bhandari, Meyer & Pardos, Proceedings of the 18th International Conference on Educational Data Mining (EDM 2025), Palermo, July 12, 2025
  2. 02GSM-Symbolic built symbolic templates over the GSM8K grade-school maths benchmark so the same problem can be re-instantiated with different names and numbers. Models show noticeable variance across instantiations of the same question; performance declines when only the numerical values change; performance deteriorates significantly as the number of clauses increases; adding a single clause that appears relevant but contributes nothing to the reasoning chain drops performance by up to 65% across all state-of-the-art models tested. The authors conclude that current models replicate reasoning steps seen in training rather than performing genuine logical reasoning. — Mirzadeh, Alizadeh, Shahrokhi, Tuzel, Bengio & Farajtabar (Apple, Washington State University), ICLR 2025, January 1, 2025
  3. 03Field experiment with nearly 1,000 high school mathematics students in Turkey, three arms: a standard chatbot interface, a tutor with prompts designed to safeguard learning, and a textbook-and-notes control. During AI-assisted practice the standard interface improved performance 48% and the safeguarded tutor 127% against control; on the exam with AI access removed, standard-interface students scored 17% worse than students who never had access, and safeguarded-tutor students scored about the same as control. — Knowledge at Wharton, summarising Bastani, Bastani, Sungu, Ge, Kabakcı & Mariman (published in PNAS), August 27, 2024
  4. 04A ten-year panel of 3.2 million ALEKS learning interactions plus 12.2 million placement-assessment response times, using a quasi-experimental design contrasting AI-susceptible text word problems with AI-resistant interactive graph problems. Learning time on AI-susceptible problems fell 2.8% per quarter among college students after ChatGPT's release, cumulating to −26.9% over eleven quarters; high schoolers −31.3%; middle schoolers −9.0%; grade 5 no detectable change. On randomly assigned proctored retention items, odds of a correct response fell 25% cumulatively. The post-ChatGPT divergence vanishes entirely under proctoring. — Rismanchian, Uzun, Matayoshi, Cosyn & Kurd-Misto, UC Irvine and McGraw Hill (preprint, under review), June 1, 2026
  5. 05Multiple GPT- and Llama-family models evaluated out of the box against human essay grades: agreement between model and human scores remains relatively weak and varies with essay characteristics; models systematically over-score short, underdeveloped essays and under-score longer essays containing minor grammatical or spelling errors; model scores are internally consistent with the model's own generated feedback — Mathew, Taher, Kundu & Barbosa (preprint), March 24, 2026
  6. 06EdTech Hub policy brief on AI in education in MENA: AI use is concentrated among urban elites, private schools and well-resourced universities, with public-sector schools less likely to adopt due to infrastructure gaps, higher student-teacher ratios and limited ministry-led integration; ministries in countries including Jordan, Iraq and Egypt have not yet embedded AI-focused training into national professional development; the brief contains no learning-outcomes evaluation — Amu & Schmitt, EdTech Hub Policy Brief, commissioned by FCDO MENA Regional Advisors, January 1, 2026
  7. 07Meta-analysis of 107 effect sizes across 14,321 participants: intelligent tutoring systems produced g = 0.42 against teacher-led large-group instruction, g = 0.35 against textbooks, and g = −0.11 (not significant) against individualised human tutoring — Ma, Adesope, Nesbit & Liu, Journal of Educational Psychology (American Psychological Association), November 1, 2014