Why the checklist focuses on independence
Two trials point in different directions because they tested different tools in different settings.
Structured tutor, positive result
A 2025 crossover trial included 194 eligible Harvard physics students. Its purpose-built tutor used expert scaffolding, targeted feedback and self-pacing. Students learned more in less time than in active-learning class sessions.
Answer machine, later harm
In a separate field experiment with nearly 1,000 high-school mathematics students in Turkey, an answer-giving GPT-4 tool raised practice performance. After the AI was removed, those students scored about 17% lower than the no-AI group. Teacher-designed guardrails largely mitigated the harm, but did not produce a positive unaided-test effect.
Do not treat these as a head-to-head product comparison. They involved different ages, subjects, designs and tutoring systems. Together they make one evaluation question more useful: what can the learner do when the tool is no longer helping?
Hear the complete evidence review
We finished this 17:17 Evidence Verdict before making the checklist. It weighs positive, negative and synthesis evidence, names the limits of each study and excludes a retracted meta-analysis.
Start the full episodeFree · no sign-up · 17:17
Episode details, transcript and sourcesEducational research commentary, not advice for a specific learner or a product endorsement. Published 2026-07-30.