All insights
Instrument plate showing the figure 0.37 mounted on a dimension line, with the two-sigma claim struck through beside it
The TechnologyMarch 10, 20266 min read

Retire the Two Sigma

The benchmark in every EdTech pitch deck rests on two graduate dissertations, three weeks of teaching, and a grading threshold set in the treatment group's favour. One of the two was never published.

If you have read an EdTech pitch deck in the last three years, you have read the two-sigma claim. A tutored student outperforms 98% of a conventionally taught class. Bloom proved it in 1984. AI makes one-to-one tutoring free. Therefore the market is everyone.

The arithmetic is fine. The benchmark is not.

Bloom's two sigma comes from a single essay in Educational Researcher, and that essay rests on two doctoral dissertations completed at the University of Chicago. Joanne Anania taught probability to fourth and fifth graders for three weeks. Arthur J. Burke taught cartography to eighth graders for three weeks. Burke never published his findings. Anania's paper has been cited 77 times in four decades.

That is the entire evidentiary base of the most repeated number in education technology.

What the study actually did

The three-week duration is the smallest problem. The larger one is that the comparison was never like-for-like.

Bloom's tutored students had to reach 90% mastery before moving on. The classroom students had to reach 80%. The treatment group was held to a higher standard, then tested and found to have met it. Kurt VanLehn, reviewing the field in 2011, concluded that Bloom's effect size "seems to be due mostly to holding tutees to a higher standard of mastery" — and measured human tutoring himself at roughly d = 0.79. Not quite two sigma. Not close to it.

The tests were also written by the researchers who designed the intervention. This matters more than it sounds. When Kulik and colleagues reanalysed mastery learning in 1990, they found an effect of d = 0.5 on tests the experimenters had written, and d = 0.08 on standardised tests. Same intervention. Same students. The effect lived almost entirely in the instrument.

Robert Slavin had already found the same thing in 1987: control for time on task, use a standardised test, and the median effect of mastery learning is approximately zero. In 2015, Jerrim and colleagues ran the comparison properly — more than 5,000 students, randomised — and measured d = 0.06, not statistically significant.

The real ceiling is 0.37

So what does one-to-one tutoring actually deliver?

Nickow, Oreopoulos and Quan pooled 96 randomised studies of tutoring in 2020 and found a mean effect of 0.37 standard deviations. Cohen, Kulik and Kulik reached 0.33 in 1982, working from a different sample four decades earlier. Two independent syntheses, separated by a generation, converge on roughly a third of a standard deviation.

For scale: most education interventions that survive contact with a randomised trial produce 0.1 SD or less. A third of a standard deviation is a genuinely good result. It is also about a sixth of what the pitch decks promise, and no study in that pool of 96 produced two sigma.

Now put an AI tutor next to it. The most careful meta-analysis in this field predates ChatGPT by eight years: Ma, Adesope, Nesbit and Liu pooled 107 effect sizes across 14,321 participants in 2014. Intelligent tutoring systems beat large-group teacher-led instruction (g = 0.42) and beat textbooks (g = 0.35). Against individual human tutoring, they measured g = −0.11, not significant.

Read that as the good news it is. Software matched a human tutor, across 14,000 participants, more than a decade ago. That is a real finding and the sector should be prouder of it.

But it also closes the argument. If AI tutoring matches human tutoring, and human tutoring is 0.37, then a perfect AI tutor lands on 0.37. The two-sigma target was never available. Chasing it means every product in the category is measured against a number it structurally cannot hit.

The best recent result, read honestly

The strongest evidence for LLM tutoring so far is a crossover randomised trial run in Harvard's largest introductory physics course in autumn 2023, published in Scientific Reports in June 2025. Every one of the 194 students took both conditions across two weeks.

The comparison condition was not a lecture. It was in-class active learning taught by experienced instructors — the strongest classroom method there is. The AI condition beat it, with an effect size of 0.63 that the authors describe as an underestimate; correcting for ceiling effects puts the range at 0.73 to 1.3 SD. Students in the AI condition also finished faster: a median of 49 minutes against roughly 60 in class.

That is a serious result and it deserves to be cited. It deserves to be cited accurately.

The tutor was built by the physics PhDs who teach the course. It ran on expert-crafted prompts with deliberate scaffolding and guardrails. The material sat at the understanding, applying and analysing levels of Bloom's taxonomy — the authors say plainly that their findings may not generalise to content requiring complex synthesis or higher-order critical thinking. The trial ran for two weeks. The authors call for replication.

The intervention here is the instructional design, not the model. Anyone can call the same API.

What this means in Arabic

Every number above was measured in English.

Anania's fourth graders, the Harvard physics students, all 96 tutoring trials, all 107 effect sizes in the 2014 meta-analysis — English-language instruction, overwhelmingly in the United States. We have not found a single randomised trial of AI tutoring conducted in Arabic.

That is not a small caveat for a fund investing in this region. It means a founder in Amman or Riyadh citing any of these effect sizes is borrowing evidence from a different language, a different curriculum and a different examination system, and presenting it as though it transfers. It might. Nobody has checked.

What a founder should do about it

Stop citing two sigma. It is the fastest way to tell an informed investor that you have not read the source. The claim has been publicly dismantled since at least 2024, and the analysis is one search away.

Cite 0.37 instead, and cite it as the ceiling. A product that reliably delivers a third of a standard deviation at a fraction of the cost of a human tutor is a good business. It is a better pitch than an impossible one, because it survives diligence.

Treat instructional design as the product. The Harvard result came from scaffolding built by people who teach the subject. If your differentiation is a system prompt over a frontier model, you have built a feature with a competitor's cost structure.

Measure in the language you sell in. The evidence gap in Arabic is a liability for anyone citing English effect sizes, and an opening for anyone willing to run the trial. A published randomised result in Arabic would be the first of its kind, and it would be worth more than any funding announcement.

What we take from it

Fikr backs seed-stage EdTech in a market where almost nothing is measured. The two-sigma claim is a useful diagnostic precisely because it is so widely repeated: it tells us, quickly, whether a founder has read the literature or absorbed the pitch.

The honest version of the case for AI in learning is smaller than the marketed one and more durable. Software can match a human tutor. Human tutoring moves outcomes by about a third of a standard deviation. Delivering that at scale, in Arabic, to students who currently get nothing like it, is a large enough problem to build a company on.

It does not need a number from 1984 that was never true.

Sources

  1. 01Bloom's 1984 essay rested on two University of Chicago dissertations; Burke's was never published; Anania's has been cited 77 times — Education Next (Harvard Kennedy School / Hoover Institution), March 1, 2024
  2. 02Bloom's tutored condition used a 90% mastery threshold against 80% for the classroom condition; VanLehn (2011) measured human tutoring at d ≈ 0.79 — Nintil (independent systematic review, not peer-reviewed), July 28, 2019
  3. 03Human tutoring produces 0.37 SD across 96 randomised studies — Nickow, Oreopoulos & Quan, NBER, July 1, 2020
  4. 04Cohen, Kulik & Kulik (1982) measured tutoring at 0.33 SD — Education Next, March 1, 2024
  5. 05Intelligent tutoring systems are statistically indistinguishable from individual human tutoring (g = −0.11, not significant); 107 effect sizes, 14,321 participants — Ma, Adesope, Nesbit & Liu, Journal of Educational Psychology (APA), November 1, 2014
  6. 06Jerrim et al. RCT with 5,000+ students found d = 0.06, not significant; Kulik et al. (1990) found d = 0.5 on experimenter-made tests and d = 0.08 on standardised tests — Nintil, July 28, 2019
  7. 07Harvard physics crossover RCT, N = 194: effect size 0.63, rising to 0.73–1.3 SD after ceiling correction; median time on task 49 minutes vs ~60 in class — Kestin et al., Scientific Reports (Nature Portfolio), June 3, 2025