# Does AI Tutoring Actually Help Students Learn? One Paper a Day, Evidence Verdict. Today is different from our usual episode. Instead of reviewing one study, we take one question people are actually asking and answer it from the weight of the evidence: does using AI as a tutor really help students learn? This is independent educational commentary about research, not advice for a specific student, teacher, or product. Every study here describes groups of people in particular settings, not a promise about your child or your classroom. I will state that boundary once, here, and then trust you to carry it. ## 1. The question and the verdict Here is the verdict first, so you have it even if you stop after this minute. A well-designed AI tutor, one with guardrails that make it coach rather than answer, can help students learn as much as, and sometimes more than, an ordinary class. But the very same technology, handed over as a chatbot that simply gives answers, can make students do worse once it is taken away, while leaving them feeling more capable, not less. The design of the tutor, not the mere presence of artificial intelligence, decides whether learning goes up or down. The biggest caveat, in the same breath: almost all of the strong evidence measures learning immediately, under supervision. Whether these gains last for months, transfer to new problems, or survive a student alone at home with a general chatbot, is largely unknown. So the honest headline is not "AI tutors work" and not "AI makes students lazy." It is "it depends, and it depends on things you can actually check." Who should care? Parents deciding whether to trust a tutoring app, teachers deciding how to let students use these tools, and anyone building or buying education technology. If you were hoping for a simple yes or no, the useful answer is more valuable than that, because it tells you what to look for. ## 2. Why this question matters now AI tutors are being deployed to millions of students right now, sold on a genuinely old and powerful promise: personal tutoring for everyone. The marketing often rests on a single viral study or a glossy demo. That is exactly the situation where a careful look at the whole body of evidence earns its keep, because individual studies point in different directions, and the differences are not random. They track the design of the tutor and the way it is used. There is also a deeper reason. Learning and performing are not the same thing. A tool can make you look better today while teaching you less for tomorrow. That gap sits at the center of this whole question, and it is where a lot of intuition goes wrong. ## 3. The field so far To judge whether today's AI tutors are a breakthrough, you have to know the bar. And the bar is old. Start with the aspiration. In 1984, the educational psychologist Benjamin Bloom described what became known as the two-sigma problem. Students tutored one-to-one, he reported, performed about two standard deviations above students in an ordinary class. Two standard deviations would move an average student near the top of the room. That number has echoed through education for forty years. But treat it as an aspiration, not a settled fact: it came from small studies and has never robustly replicated. It tells us tutoring is powerful. It does not tell us to expect a doubling from any tool that calls itself a tutor. Now the reality check. In 2011, Kurt VanLehn reviewed decades of experiments comparing human tutors, computer tutors, and no tutoring. His finding reset expectations twice over. Human tutoring was not the mythical two sigma; it landed closer to an effect size of about zero point seven nine. And the older intelligent tutoring systems, software built long before large language models, reached about zero point seven six. In other words, well-built computer tutors were already roughly as effective as human tutors, and both delivered something closer to eight tenths of a standard deviation than to two. That is the crucial context. The question "can a computer tutor?" was answered years ago, and the answer was a qualified yes. So the real question for today's AI is narrower and harder: do large-language-model tutors clear that established bar, and do their gains actually last? ## 4. The evidence that AI tutoring helps Three lines of recent evidence say that, done well, it does. The strongest single trial comes from Harvard, published in Scientific Reports in 2025 by Gregory Kestin and colleagues. About one hundred and eighty physics students alternated between two conditions: a well-run, active-learning class, and studying the same material at home with a purpose-built AI tutor. This was not a raw chatbot. It was carefully engineered with expert-written scaffolds, step-by-step prompting, and guardrails against just handing over answers. The result: students learned more with the AI tutor, in less time, scoring roughly thirty percent higher on the test afterward, which the authors frame as about double the learning gains of the class. They also felt more engaged and motivated. That is a striking result, and it deserves its boundaries. These were motivated students at an elite university, the topics were short and well-defined, the design compared each student against themselves week to week, and the tutor was heavily refined. The finding attaches to that tutor, in that setting, not to chatbots in general. The second line widens the setting dramatically. A 2025 World Bank study in Nigeria, a working paper by Martin De Simone and colleagues, ran a six-week after-school program in Benin City where students worked with a tutor built on GPT-four. Learning on an English-focused assessment rose by about zero point two three standard deviations, and the authors describe the overall gains as comparable to roughly two years of ordinary schooling, at a cost near forty-eight dollars per student. Two boundaries matter: this is a working paper, not yet peer reviewed, and the sessions were supervised. But it shows real gains outside a wealthy university, which is where the two-sigma dream was always aimed. The third line is the aggregate. A 2026 meta-analysis pooling thirty-five experimental studies found a moderate, statistically significant positive effect of ChatGPT-style tools on learning, an effect size around zero point six seven. Other recent meta-analyses land in a similar band, roughly zero point four five to zero point eight six. But read that number honestly: it is an average across very different studies, with a lot of variation by subject, by how long the study ran, and by how the tool was used. A moderate average is not a promise that any given classroom will see the same. ## 5. The evidence that it can backfire Now the study that keeps the verdict honest, and that most of the marketing ignores. In 2025, Hamsa Bastani, Osbert Bastani, and colleagues published a randomized trial in the Proceedings of the National Academy of Sciences, working with nearly a thousand high-school students in Turkey. They compared two AI tutors. One, call it the plain chatbot, worked like standard ChatGPT: ask it, and it answers. The other was a guarded tutor, prompted to give teacher-designed hints instead of solutions. During practice, both helped, and the plain chatbot helped a lot; grades on assisted problems jumped. Then the researchers took the AI away and tested students on their own. The students who had used the plain chatbot now did worse than classmates who never had AI at all, by around seventeen percent. They had leaned on it, felt productive, and learned less. The guarded tutor avoided that harm, but here is the sobering part: it did not durably beat the students who had no AI either. This is the illusion of learning made concrete. In-the-moment performance went up while durable skill went down. And the difference between harming learning and protecting it was not the intelligence of the model. It was one design choice: hints versus answers. ## 6. What is settled and what is still open Put those together and a clear shape emerges. What is genuinely settled: a well-designed, guarded AI tutor can produce real learning gains, in the same broad band as the good tutoring systems that came before it, when used under supervision for a defined stretch. Equally settled: an unguarded, answer-giving chatbot can reduce durable learning even as it flatters performance. The guardrails are the active ingredient. What is still open, and it is a lot: whether these gains persist weeks or months later, because most trials test immediately; whether they transfer to genuinely new problems rather than near-copies; and what happens in the real default case, a student alone at home, unsupervised, with a general chatbot that has every incentive to be helpful by just answering. The newest work is starting to probe this. An exploratory 2025 trial in United Kingdom classrooms suggests guarded AI tutoring can support students safely, but it is small and preliminary, a direction to watch rather than a conclusion. ## 7. Limits and one integrity note A few honest limits across all of this. The positive trials often use purpose-built, expensive tutors, not the free apps most students actually reach for. Several strong results, including the Nigeria study, are working papers or preprints, not yet through peer review. And the headline numbers, "double the gains," "two years in six weeks," come from specific settings and should not be treated as portable guarantees. One integrity note, because it matters for how you read coverage of this topic. A widely shared meta-analysis on ChatGPT and learning was retracted. I have deliberately not used it as evidence anywhere in this episode; the synthesis I cited is a separate, non-retracted analysis. When a field is hot and the studies are small, the secondary literature gets noisy, and a retracted paper can keep circulating long after it should. Checking whether a striking claim rests on a retracted source is part of reading responsibly. ## 8. What you can take away Five takeaways you could repeat to a friend, each labeled by how settled the evidence is. One. A well-designed, guarded AI tutor beat a strong active-learning class in a Harvard trial, with about double the learning gains in less time. Well supported here, for that tutor and setting. Two. Across about thirty-five studies, ChatGPT-style tools show a moderate average boost to learning, roughly zero point six seven. Well supported as an average, but suggestive for any single classroom, because the variation is large. Three. An unguarded chatbot that gives answers can make students perform worse on their own afterward, around seventeen percent worse in one large trial, while feeling more capable. Well supported here: the illusion of learning is real. Four. The guardrails, hints instead of answers, are the active ingredient. "AI or no AI" is the wrong question; "which design" is the right one. Well supported here. Five. Whether these gains last for months, transfer to new problems, or survive unsupervised use of a general chatbot is still unknown. Still unknown, and the most important thing left to learn. Notice what these separate: what the evidence shows, what it leaves open, and what would change the verdict. Large trials with delayed tests, weeks or months later, and real unsupervised use, independently replicated, would move this from "promising with conditions" to something firmer. ## 9. What to actually do with this For a parent or a student, the practical rule follows the evidence. A tool that makes you struggle a little and gives hints is working with your learning; a tool that hands you finished answers is quietly borrowing against it. The feeling of ease is not the signal of learning, and can be the opposite. The useful test is simple: after the AI is closed, can you do the next problem yourself? For a teacher or a builder, the lesson is sharper. Do not buy or sell "AI tutoring" as a category; the category contains both the Harvard result and the Turkey harm. Ask what the tool does when a student is stuck: does it coach, or does it complete? Ask whether the evidence behind it measured learning after the tool was removed, not just during use. And treat the design of the guardrails, not the size of the model, as the thing that determines whether students end up knowing more. ## 10. Skeptical checklist and exact source card Six quick questions before trusting any AI-tutoring claim. Does the tutor give hints or answers? Was learning measured after the tool was taken away, or only while it was in hand? How long afterward? Was it supervised or self-directed? Is the study peer reviewed or a preprint? And is a product being sold on the back of it? The sources, each cited from its own record. The strongest confirming trial is "AI tutoring outperforms in-class active learning," by Gregory Kestin and colleagues, in Scientific Reports, 2025; the digital object identifier is ten point one oh three eight slash s four one five nine eight dash oh two five dash nine seven six five two dash six. The field-deployment evidence is "From Chalkboards to Chatbots," by Martin De Simone and colleagues, World Bank Policy Research Working Paper eleven twenty-five, 2025, a working paper. The conflicting evidence is "Generative AI without guardrails can harm learning," by Hamsa Bastani and colleagues, in the Proceedings of the National Academy of Sciences, 2025; the identifier is ten point one oh seven three slash pnas dot two four two two six three three one two two. The synthesis is a 2026 meta-analysis of thirty-five studies in Humanities and Social Sciences Communications. The foundation is Benjamin Bloom's 1984 two-sigma paper in Educational Researcher, and Kurt VanLehn's 2011 review in Educational Psychologist. The most recent signal is an exploratory 2025 United Kingdom classroom trial, arXiv two five one two point two three six three three, a preprint. One meta-analysis on this topic has been retracted and was deliberately excluded from this episode. This is newly written commentary and does not reproduce any paper's prose, figures, or tables. The final verdict: AI tutoring genuinely helps when it is built to coach, and can genuinely hurt when it is built to answer, and today the strongest evidence measures the moment of learning rather than what lasts. Judge the tool by what it does when a student is stuck, and by whether the learning survives the tool being switched off. That, not the word "AI," is where the answer lives.