The strongest result in AI and education is not a chatbot talking to a student. It is a chatbot talking to a tutor.
Stanford researchers randomised 900 tutors serving 1,800 K-12 students from historically under-served communities, with a preregistered analysis plan, and gave half the tutors a real-time AI assistant during live sessions. Students whose tutor had it were 4 percentage points more likely to master the topic. Among students working with the lowest-rated tutors, mastery rose by 9 percentage points.
The cost, based on actual usage, was $20 per tutor per year.
That combination — a measured gain, concentrated where the human capital is weakest, at a price a school district would not notice — has no equivalent anywhere else in this literature. And it points at a product category most founders are not building.
Why replacing the tutor was never the play
The ceiling for a student-facing AI tutor was established before ChatGPT existed, and almost nobody in the sector quotes it.
Ma, Adesope, Nesbit and Liu pooled 107 effect sizes across 14,321 participants in 2014. Intelligent tutoring systems beat large-group teacher-led instruction at g = 0.42 and beat textbooks at g = 0.35. Against individualised human tutoring, they measured g = −0.11, not statistically significant.
That is a genuinely good finding and the sector should cite it more. Software matched a human tutor, across 14,321 participants, more than a decade ago. It also sets the ceiling: a perfect AI tutor matches a human tutor. It does not exceed one. Every product aimed at replacing the tutor is competing for a tie.
Coaching the tutor is a different problem with a different ceiling, because it moves the human, and the human was the thing that worked.
What the student-facing evidence actually shows
The best-known disconfirming study in the field ran with nearly 1,000 high school mathematics students in Turkey, in three arms: a standard chatbot interface, a tutor built with prompts designed to safeguard learning, and a control with textbook and notes only.
During AI-assisted practice, the standard interface improved performance by 48% against control and the safeguarded tutor by 127%. Then the researchers took the AI away and ran an exam.
Students who had used the standard interface scored 17% worse than students who never had access at all. Students who had used the safeguarded tutor scored about the same as control.
Read the second line as carefully as the first. The guardrailed tutor did not produce a learning gain. It produced no harm. The 127% improvement during practice was performance, not learning — and it was the larger of the two in-product numbers. Any metric derived from what a student does inside your product is measuring the wrong thing, and measuring it most flatteringly in the arm that performed worst on the exam.
One caution on those percentages. They come from Wharton's own summary of the working paper, dated August 2024; the published version sits behind a paywall we could not open to confirm the figures match. The direction of the finding is not in dispute; the exact percentages should be verified against the journal version before they go in a deck.
The one recent student-facing result that looks strong turns out, on inspection, to support the same conclusion. A 2025 exploratory trial in five UK secondary schools reported that students supported by Google's LearnLM were 5.5 percentage points more likely to solve novel problems on later topics — 66.2% against 60.7% for human tutoring alone. But expert human tutors supervised the model, with the remit to revise every drafted message until they would happily send it themselves. They approved 76.4% of drafted messages with no or minimal edits. That is not autonomous AI tutoring. It is a human tutor with a drafting assistant, which is the same intervention as Tutor CoPilot approached from the other end. It is also vendor research — run by the model's developer with a commercial platform partner, on 165 students, and described by its own authors as exploratory.
The mechanism is legible, which is rare
Most education technology asks you to believe that an outcome happened for the reason the vendor says. Tutor CoPilot does not.
The researchers classified more than 550,000 messages from live sessions. Tutors with the assistant asked more guiding questions and gave less generic praise. The pedagogy improved in the transcript, not in a survey, and the improvement is the kind an education researcher recognises as real teaching behaviour rather than engagement.
That is also why the effect concentrates where it does. The assistant is transferring the practice of a good tutor to a tutor who did not have it. A strong tutor already asks guiding questions, so there is little to add. A weak one gains 9 points for their students.
The honest caveats belong here too. Tutor CoPilot is a preprint and working paper, not a peer-reviewed publication. It ran with one tutoring provider in the United States, in English, with K-12 students. The overall effect of 4 percentage points is real but not large. And there has been no replication in another language, another country, or another delivery model.
What this means for the region
The constraint in MENA education is not that students lack a chatbot. It is that the adults in the room have no support.
EdTech Hub's regional brief, commissioned by FCDO advisors and published in January 2026, records that ministries in Jordan, Iraq and Egypt have not embedded AI-focused training into national professional development. Where teachers are using these tools, they are doing it without institutional guidance.
The nearest measurement of that gap is American, and it is unflattering. Gallup finds 60% of US public school teachers used an AI tool for work during the school year and 32% use one at least weekly, while only 18% receive any formal guidance on AI use, and 69% receive none at all on using AI for one-on-one instruction or tutoring. Do not extrapolate those percentages to Amman or Cairo; no comparable regional survey exists. Do extrapolate the shape of the problem, because the EdTech Hub brief describes the same shape qualitatively.
There is a second signal in how educators already behave when given the tools. Anthropic's analysis of roughly 74,000 conversations from higher-education professionals found curriculum development at 57% of use, academic research at 13%, and assessing student performance at 7%. Teaching and instruction ran 77.4% augmentation. Grading and student records ran 48.9% automation. That is vendor research on one product's users and it observes no outcome. It does describe what educators reach for unprompted, and it is preparation and design, not delegation of the teaching itself.
[FIKR TO CONFIRM: whether the fund wants to state a formal thesis preference for educator-facing over student-facing AI products, and whether any portfolio company already sits in this category.]
What a founder should do about it
Point the model at the adult. The measured evidence favours products that make an average educator better over products that replace one. That is a smaller-sounding ambition and a much better business, because the buyer already exists, the distribution already exists, and the outcome is attributable.
Price like the result. $20 per tutor per year is the number in the study. It suggests the winning shape is a low per-seat price against a large seat count inside institutions, not a consumer subscription fighting for a parent's attention.
Instrument the transcript. Tutor CoPilot's most persuasive evidence is a classification of 550,000 real messages showing pedagogy changed. If your product sits in a conversation, you already have the corpus to run the same analysis. Almost nobody does, and it is the strongest diligence artefact you could bring to a first meeting.
Never report an in-product metric as a learning outcome. The Turkish trial's 127% practice improvement came with zero exam gain. Report the delta on an assessment the student takes without your product in the room, or say plainly that you have not measured one.
Run the replication nobody has run. There is no Arabic-language trial of AI coaching for tutors or teachers. At $20 per tutor per year, this is among the cheapest category-defining pieces of evidence available to anyone in this region, and the first credible result would be cited for years.
The pitch that sells is the AI tutor that teaches every child. The evidence supports something less cinematic: an AI that makes a mediocre human educator noticeably better, for twenty dollars a year, with the largest gain going to the students who currently have the weakest teacher. That is the category with the numbers behind it.




