All insights
Instrument plate showing a 9-point interval on a dimension line, with a tutor marker and a student marker positioned either side of it
The TechnologyApril 29, 20267 min read

The Best AI-in-Education Result Is Not a Tutor

The largest measured gain in the field came from putting AI behind 900 human tutors rather than in front of students. It cost $20 per tutor per year, and it worked hardest for the students of the weakest tutors.

The strongest result in AI and education is not a chatbot talking to a student. It is a chatbot talking to a tutor.

Stanford researchers randomised 900 tutors serving 1,800 K-12 students from historically under-served communities, with a preregistered analysis plan, and gave half the tutors a real-time AI assistant during live sessions. Students whose tutor had it were 4 percentage points more likely to master the topic. Among students working with the lowest-rated tutors, mastery rose by 9 percentage points.

The cost, based on actual usage, was $20 per tutor per year.

That combination — a measured gain, concentrated where the human capital is weakest, at a price a school district would not notice — has no equivalent anywhere else in this literature. And it points at a product category most founders are not building.

Why replacing the tutor was never the play

The ceiling for a student-facing AI tutor was established before ChatGPT existed, and almost nobody in the sector quotes it.

Ma, Adesope, Nesbit and Liu pooled 107 effect sizes across 14,321 participants in 2014. Intelligent tutoring systems beat large-group teacher-led instruction at g = 0.42 and beat textbooks at g = 0.35. Against individualised human tutoring, they measured g = −0.11, not statistically significant.

That is a genuinely good finding and the sector should cite it more. Software matched a human tutor, across 14,321 participants, more than a decade ago. It also sets the ceiling: a perfect AI tutor matches a human tutor. It does not exceed one. Every product aimed at replacing the tutor is competing for a tie.

Coaching the tutor is a different problem with a different ceiling, because it moves the human, and the human was the thing that worked.

What the student-facing evidence actually shows

The best-known disconfirming study in the field ran with nearly 1,000 high school mathematics students in Turkey, in three arms: a standard chatbot interface, a tutor built with prompts designed to safeguard learning, and a control with textbook and notes only.

During AI-assisted practice, the standard interface improved performance by 48% against control and the safeguarded tutor by 127%. Then the researchers took the AI away and ran an exam.

Students who had used the standard interface scored 17% worse than students who never had access at all. Students who had used the safeguarded tutor scored about the same as control.

Read the second line as carefully as the first. The guardrailed tutor did not produce a learning gain. It produced no harm. The 127% improvement during practice was performance, not learning — and it was the larger of the two in-product numbers. Any metric derived from what a student does inside your product is measuring the wrong thing, and measuring it most flatteringly in the arm that performed worst on the exam.

One caution on those percentages. They come from Wharton's own summary of the working paper, dated August 2024; the published version sits behind a paywall we could not open to confirm the figures match. The direction of the finding is not in dispute; the exact percentages should be verified against the journal version before they go in a deck.

The one recent student-facing result that looks strong turns out, on inspection, to support the same conclusion. A 2025 exploratory trial in five UK secondary schools reported that students supported by Google's LearnLM were 5.5 percentage points more likely to solve novel problems on later topics — 66.2% against 60.7% for human tutoring alone. But expert human tutors supervised the model, with the remit to revise every drafted message until they would happily send it themselves. They approved 76.4% of drafted messages with no or minimal edits. That is not autonomous AI tutoring. It is a human tutor with a drafting assistant, which is the same intervention as Tutor CoPilot approached from the other end. It is also vendor research — run by the model's developer with a commercial platform partner, on 165 students, and described by its own authors as exploratory.

The mechanism is legible, which is rare

Most education technology asks you to believe that an outcome happened for the reason the vendor says. Tutor CoPilot does not.

The researchers classified more than 550,000 messages from live sessions. Tutors with the assistant asked more guiding questions and gave less generic praise. The pedagogy improved in the transcript, not in a survey, and the improvement is the kind an education researcher recognises as real teaching behaviour rather than engagement.

That is also why the effect concentrates where it does. The assistant is transferring the practice of a good tutor to a tutor who did not have it. A strong tutor already asks guiding questions, so there is little to add. A weak one gains 9 points for their students.

The honest caveats belong here too. Tutor CoPilot is a preprint and working paper, not a peer-reviewed publication. It ran with one tutoring provider in the United States, in English, with K-12 students. The overall effect of 4 percentage points is real but not large. And there has been no replication in another language, another country, or another delivery model.

What this means for the region

The constraint in MENA education is not that students lack a chatbot. It is that the adults in the room have no support.

EdTech Hub's regional brief, commissioned by FCDO advisors and published in January 2026, records that ministries in Jordan, Iraq and Egypt have not embedded AI-focused training into national professional development. Where teachers are using these tools, they are doing it without institutional guidance.

The nearest measurement of that gap is American, and it is unflattering. Gallup finds 60% of US public school teachers used an AI tool for work during the school year and 32% use one at least weekly, while only 18% receive any formal guidance on AI use, and 69% receive none at all on using AI for one-on-one instruction or tutoring. Do not extrapolate those percentages to Amman or Cairo; no comparable regional survey exists. Do extrapolate the shape of the problem, because the EdTech Hub brief describes the same shape qualitatively.

There is a second signal in how educators already behave when given the tools. Anthropic's analysis of roughly 74,000 conversations from higher-education professionals found curriculum development at 57% of use, academic research at 13%, and assessing student performance at 7%. Teaching and instruction ran 77.4% augmentation. Grading and student records ran 48.9% automation. That is vendor research on one product's users and it observes no outcome. It does describe what educators reach for unprompted, and it is preparation and design, not delegation of the teaching itself.

[FIKR TO CONFIRM: whether the fund wants to state a formal thesis preference for educator-facing over student-facing AI products, and whether any portfolio company already sits in this category.]

What a founder should do about it

Point the model at the adult. The measured evidence favours products that make an average educator better over products that replace one. That is a smaller-sounding ambition and a much better business, because the buyer already exists, the distribution already exists, and the outcome is attributable.

Price like the result. $20 per tutor per year is the number in the study. It suggests the winning shape is a low per-seat price against a large seat count inside institutions, not a consumer subscription fighting for a parent's attention.

Instrument the transcript. Tutor CoPilot's most persuasive evidence is a classification of 550,000 real messages showing pedagogy changed. If your product sits in a conversation, you already have the corpus to run the same analysis. Almost nobody does, and it is the strongest diligence artefact you could bring to a first meeting.

Never report an in-product metric as a learning outcome. The Turkish trial's 127% practice improvement came with zero exam gain. Report the delta on an assessment the student takes without your product in the room, or say plainly that you have not measured one.

Run the replication nobody has run. There is no Arabic-language trial of AI coaching for tutors or teachers. At $20 per tutor per year, this is among the cheapest category-defining pieces of evidence available to anyone in this region, and the first credible result would be cited for years.

The pitch that sells is the AI tutor that teaches every child. The evidence supports something less cinematic: an AI that makes a mediocre human educator noticeably better, for twenty dollars a year, with the largest gain going to the students who currently have the weakest teacher. That is the category with the numbers behind it.

Sources

  1. 01Tutor CoPilot randomised 900 tutors serving 1,800 K-12 students from historically under-served communities, with a preregistered analysis plan; students of tutors given access were 4 percentage points more likely to master topics (p < 0.01), and students of lower-rated tutors improved mastery by 9 percentage points relative to control; cost was $20 per tutor per year based on actual usage; over 550,000 messages were classified, showing tutors with the tool asked more guiding questions and gave less generic praise — Wang, Ribeiro, Robinson, Loeb & Demszky, Stanford University (preprint / working paper), October 1, 2024
  2. 02Field experiment with nearly 1,000 high school mathematics students in Turkey, three arms: standard chatbot interface, a tutor with prompts designed to safeguard learning, and a textbook-and-notes control; during AI-assisted practice the standard interface improved performance 48% and the safeguarded tutor 127% against control; on the exam with AI access removed, standard-interface students scored 17% worse than students who never had access, and safeguarded-tutor students scored about the same as control — Knowledge at Wharton, summarising Bastani et al. (working paper; published in PNAS), August 27, 2024
  3. 03Ma, Adesope, Nesbit & Liu pooled 107 effect sizes across 14,321 participants and found intelligent tutoring systems beat teacher-led large-group instruction at g = 0.42 and textbooks at g = 0.35, but were statistically indistinguishable from individualised human tutoring at g = −0.11, not significant — Ma, Adesope, Nesbit & Liu, Journal of Educational Psychology (APA), November 1, 2014
  4. 04Exploratory RCT with 165 students across five UK secondary schools in mathematics, in which expert human tutors supervised and revised every drafted AI message; tutors approved 76.4% of drafted messages with zero or minimal edits, and students supported by the system were 5.5 percentage points more likely to solve novel problems on subsequent topics, 66.2% against 60.7% for human tutoring alone; the study was run by the model's developer with a commercial platform partner — LearnLM Team (Google) & Eedi (preprint, vendor research), December 29, 2025
  5. 05Analysis of approximately 74,000 anonymised conversations from higher-education professionals plus interviews with 22 Northeastern University faculty: curriculum development 57% of use, academic research 13%, assessing student performance 7%; teaching and instruction ran 77.4% augmentation, while grading and student records ran 48.9% automation — Anthropic Education Report: How Educators Use Claude (vendor research), August 27, 2025
  6. 06Gallup survey of 2,069 US public K-12 teachers finds 18% receive any formal guidance on AI tool use, 34% receive no guidance at all across the ten measured tasks, and 69% get no guidance on using AI for one-on-one instruction or tutoring — Gallup, May 26, 2026
  7. 07Gallup / Walton Family Foundation survey of 2,232 US public K-12 teachers finds 60% used an AI tool for work during the school year and 32% use AI at least weekly — Gallup / Walton Family Foundation, June 24, 2025
  8. 08EdTech Hub's MENA policy brief records that ministries in countries such as Jordan, Iraq and Egypt have not yet embedded AI-focused training into national professional development, and contains no learning-outcomes evaluation of AI in education anywhere in the region — Amu & Schmitt, EdTech Hub Policy Brief, commissioned by FCDO MENA Regional Advisors, January 1, 2026