The Tutor Problem

Photo: Unsplash

AI Tutoring

The Tutor Problem

Bloom's two-sigma effect is real. Whether AI delivers it depends entirely on which half of tutoring you're replacing.
AI-tutoringeducationlearning-scienceEdTechcognition

Benjamin Bloom published a paper in 1984 called “The 2 Sigma Problem”, and the finding is simple enough to fit in a sentence: students given one-to-one tutoring performed about two standard deviations better than students in ordinary classroom instruction. A student sitting at the median in a classroom lands somewhere near the 98th percentile with a tutor.

The problem in the title is that one-to-one human tutoring cannot be provided at scale for any price a school system can pay.

Solving that is the entire promise of AI tutoring, and it’s a worthy goal. Whether the current products are achieving it is a different question, and the answer turns on a distinction almost nobody in the market makes.

Several very different things get filed under “AI tutor”. Systems that provide worked examples and check whether the student can reproduce them. Systems that identify a gap and select from a library of prepared explanations. Systems that use a language model to hold an open-ended conversation about the material. And systems claiming all three.

The first two have decades of evidence behind them. The third is where the money is going, and the evidence is thin.

Intelligent Tutoring Systems, the category that predates any of this, have been studied rigorously for forty years. Carnegie Learning’s MATHia, which grew out of the Cognitive Tutor work at Carnegie Mellon in the 1990s, probably has the strongest evidence base of any educational software ever shipped. Independent evaluations found real gains against control groups, and meta-analyses across dozens of ITS studies land on effect sizes around half a standard deviation. Not two sigma. Substantial.

Those systems work by modelling the student’s knowledge state explicitly, selecting problems calibrated to that state, and giving feedback at the level of the individual step rather than restating the general concept.

Current language-model tutors mostly don’t do any of that. They answer questions conversationally, which is useful, but they don’t maintain a rigorous model of what this particular student knows, don’t systematically choose the next problem to build a specific skill, and don’t diagnose errors at step level.

The gap shows up immediately in practice. Take a student learning to solve quadratics. A knowledge-modelling system tracks every step, identifies the specific misconception producing the error, say a sign flip when completing the square, and teaches against that misconception before moving on. A conversational tutor handles “I got -3 and -5, is that right?” by checking the answer and explaining if it’s wrong.

The second is better than nothing. It is dramatically worse than the first, and the reason is architectural rather than a matter of model quality.

Language models don’t maintain persistent state about a learner the way a purpose-built ITS does. Each conversation is largely new. History can be stuffed into a context window, but reading a transcript is not the same operation as running an explicit student model. These systems are trained to produce plausible responses, not to maintain rigorous cognitive models of individuals. Different tasks, different architectures, and no amount of scale collapses the difference.

The companies building this know it, and the honest ones will say so off the record: language models are good at explaining a concept four different ways, answering in natural language, and drawing connections across topics. They’re weak at precise knowledge modelling, calibrated problem selection and step-level diagnosis. The two technologies are complementary rather than competing.

That framing is correct and it is almost entirely absent from investment decks, which collapse everything into a single category because a single category is easier to price.

The systems that actually work combine both: an ITS core doing the modelling, with a language layer handling the explanatory conversation. It’s harder to build than either alone, because it requires stitching together two paradigms developed by communities that have historically not talked to each other. It’s also where the real progress is.

There’s a deeper question underneath all of this that cognitive scientists still argue about, and it deserves surfacing because it bounds everything.

Bloom’s paper assumes the mechanism behind the effect is adaptation and feedback. A competing account says a large slice of it is motivational: the student does better because a human being is watching who cares whether they succeed, because there’s social pressure not to quit, because the tutor’s attention is itself reinforcing.

If the motivational share is large, AI tutors face a ceiling that no technical improvement will lift, because the active ingredient is relationship rather than instructional customisation.

The evidence genuinely cuts both ways. ITS systems provide no relational component and still produce significant gains, which says the adaptation mechanism matters on its own. But human tutors consistently outperform ITS systems, and the gap is widest for struggling students, which is exactly the pattern you’d predict if motivation matters most when baseline engagement is lowest.

So: these systems are well placed to replace the adaptive-instruction half of tutoring and poorly placed to replace the other half. For a student who’s already motivated and mainly needs better explanation and better-calibrated practice, an AI tutor gets genuinely close to the human equivalent. For a student who’s disengaged, who doesn’t believe they can do it, who needs someone to notice they’re in the room, it offers considerably less.

Which points at a deployment strategy that most of the sector is ignoring.

Target these systems at students who are motivated but under-resourced. Students who want to learn and have no access to a tutor, who can’t afford test prep, who need explanation on demand in a subject where their school has nobody strong. For that population this is a real intervention with real upside, and the case is easy to make.

Deploying the same thing as a general replacement for human educational support, which several large districts have drifted toward mostly for cost reasons, is a different proposition with much weaker evidence behind it. The students who most need human engagement are precisely the ones most likely to be harmed by swapping human support for software.

That’s not an argument against AI tutoring. It’s an argument for deploying it where the evidence supports it and not where it doesn’t.

One more split worth making: content tutoring versus metacognitive tutoring.

Content tutoring is explaining photosynthesis, walking through long division, clarifying what actually started the First World War. This is where current systems are strongest, and the strength is real. Explanation quality is high, coverage is broad, and being able to ask a fourth follow-up question until something finally lands beats both a static textbook and a teacher with thirty students all wanting attention at once.

Metacognitive tutoring is teaching a student to know what they don’t know, to monitor their own comprehension, to notice when an explanation has left a gap. That’s much harder, and it’s the exact skill that separates students who use these tools well from students who use them as an answer service.

An AI that detects confusion and explains the concept again is doing content tutoring. An AI that notices the student is moving too fast, confirming understanding at the surface and missing the structure underneath, and interrupts to make them explain it back, would be doing metacognitive coaching. Very few systems do the second reliably, and I’m not aware of one that does it well.

Which produces an uncomfortable circularity, and it’s the honest conclusion of the whole thing. The students who get the most from these tools are the ones who already have the metacognitive skills: they know when they don’t understand, they push back on an unsatisfying explanation, they ask the fourth question. The students who get the least are the ones who accept whatever appears, move on when told they’re correct, and never develop the internal check that would have caught the shallowness.

The tool amplifies what the student brings to it. And the students who most need a tutor, who most need somebody to model curiosity and scepticism until they can run it themselves, are the ones who get least from a system deployed without that human scaffolding.

So the sequence matters more than the technology. A school that runs AI tutoring alongside instruction explicitly aimed at building those habits will see different outcomes from a school that runs the same software as a way to reduce human instructional hours. Same product, opposite result. Bloom’s two sigma is real and reachable. The route runs through the human part first.

Get the next live webinar in your inbox

One email a month: the upcoming live event + free recording access for subscribers. No spam, unsubscribe anytime.