---
title: "What the evidence says about AI tutors | DuckyHelper"
url: "https://duckyhelper.com/research/principles/ai-tutoring-evidence/"
description: "Randomized trials of AI tutors so far: Harvard physics, Ghana, Nigeria, Tutor CoPilot, and the PNAS study where AI help hurt. What they show, and do not."
updated: "2026-10-09"
---

AI tutors

# What the evidence says about AI tutors

Early trials are promising when the AI is built to teach: a Harvard physics tutor, a math tutor in Ghana and an English program in Nigeria all found learning gains. When students got a plain chatbot, a 2025 PNAS study found exam grades fell once it was gone. The design decides it. DuckyHelper is built to teach, and has not itself been tested in a trial yet.

> “students learn significantly more in less time when using the AI tutor”

[Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15(1), 17458.](https://doi.org/10.1038/s41598-025-97652-6)

From Kestin et al. (2025), Scientific Reports. Reproduced for commentary. Source: [Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15(1), 17458.](https://doi.org/10.1038/s41598-025-97652-6).

In DuckyHelper

## Built like the tutors that helped

In the film, a learner asks for the answer and the PNAS page shows the cost of a plain chatbot. DuckyHelper follows the designs that worked in these trials:

- Built to teach, not to answer (Mac and web)
- Trained to teach (Mac and web)
- Learning science in the tools (Mac and web)

## What the research found

### When the AI is built to teach

At Harvard, Gregory Kestin and colleagues compared an AI tutor, designed with the same teaching practices as the course, against an in-class active learning lesson. Students learned significantly more, in less time, with the AI tutor (the line quoted at the top of this page).

In Ghana, about 500 students used Rori, a math tutor on WhatsApp, for one hour a week during study hall. Their growth scores were substantially higher, an effect size of 0.36 (Henkel et al., 2024). In Nigeria, a six-week after-school program with Microsoft Copilot (GPT-4) raised English results:

> The effect on English, the main outcome of interest, was of 0.23 standard deviations (De Simone et al., 2025)

AI can also help human tutors. In Tutor CoPilot, students whose tutors had AI suggestions were 4 percentage points more likely to master topics, and the tutors were

> less likely to give away the answer to the student (Wang et al., 2024b)

ChatGPT-written hints produced learning gains like human-written hints in Pardos and Bhandari (2024).

### When the AI just answers

> unfettered access to GPT-4 can harm educational outcomes (Bastani et al., 2025)

In that study the plain chatbot raised practice grades and cut exam grades by 17%, while the version with learning safeguards mostly avoided the harm. See [Guide, don't hand over answers](https://duckyhelper.com/research/principles/guide-dont-hand-over-answers/).

### Measuring teaching, not just answering

Research groups are now testing whether models teach well, not only whether they are right. Google's LearnLM work found pedagogy-tuned models preferred by educators, and an evaluation taxonomy for AI tutors separates the two jobs:

> highlighting which LLMs are good tutors and which ones are more suitable as question-answering systems (Maurya et al., 2025)

> subject expertise, indicated by solving ability, does not immediately translate to good teaching (Macina et al., 2025)

Older research on computer tutors is the long view: reviews found intelligent tutoring systems raised test scores (Kulik and Fletcher, 2016; Ma et al., 2014), nearly as much as human tutors in VanLehn (2011).

## Feature by feature

### Built to teach, not to answer

A tutor, not an answer machine: it asks, hints and checks, following the safeguards that worked in the trials above. *Mac and web.*

### Trained to teach

DuckyHelper's tutor is trained on 1,000,000+ tutoring examples and inspired by 100+ learning-science papers, the library on this site. *Mac and web.*

### Learning science in the tools

Quizzes, spaced flashcards, worked steps, hints and drawing are each tied to a principle on this site. *Mac and web.*

What this page does not claim

No study of DuckyHelper exists yet. Every number on this page belongs to the cited study and the tool it tested, not to DuckyHelper.

## The papers on this page

### AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting

Kestin et al., 2025. Scientific Reports. Students learned more in less time with an AI tutor than in class.

### Effective and Scalable Math Support: Experimental Evidence on the Impact of an AI-Math Tutor in Ghana

Henkel et al., 2024. AIED. A field experiment with about 500 students in Ghana using Rori, a WhatsApp math tutor, one hour a week during study hall.

### From Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in Nigeria

De Simone et al., 2025. World Bank Policy Research Working Paper. A randomized trial in Nigeria where secondary students used Microsoft Copilot (GPT-4) for English as an after-school tutor for six weeks.

### Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise

Wang et al., 2024. arXiv preprint (v1). AI guidance for human tutors raised student topic mastery by 4 points.

### ChatGPT-generated help produces learning gains equivalent to human tutor-authored help on mathematics skills

Pardos & Bhandari, 2024. PLOS ONE. ChatGPT-written hints produced learning gains comparable to human tutor hints.

### Generative AI without guardrails can harm learning: Evidence from high school mathematics

Bastani et al., 2025. Proceedings of the National Academy of Sciences (PNAS). Unguarded GPT-4 help lifted practice grades, then cut exam grades 17%.

### Towards Responsible Development of Generative AI for Education: An Evaluation-Driven Approach

Jurenka et al., 2024. arXiv (Google DeepMind tech report). Educators and learners preferred a pedagogy-tuned Gemini tutor over prompting alone.

### LearnLM: Improving Gemini for Learning

LearnLM Team, Google, 2024. arXiv (Google tech report). Training on teaching instructions made Gemini the experts' pick for learning.

### Evaluating Gemini in an arena for learning

LearnLM Team, Google, 2025. arXiv (Google tech report). Educators ran blind, head-to-head tutoring comparisons of leading AI models.

### Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors

Maurya et al., 2025. NAACL 2025. An eight-dimension benchmark for judging how well LLMs actually tutor.

### MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors

Macina et al., 2025. EMNLP. An open benchmark for how well LLMs teach math, with a reward model that tells expert from novice tutor replies; solving skill did not translate into good teaching.

### The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems

VanLehn, 2011. Educational Psychologist. Step-based tutoring systems came nearly as close as human tutors.

### Effectiveness of Intelligent Tutoring Systems: A Meta-Analytic Review

Kulik & Fletcher, 2016. Review of Educational Research. A meta-analysis of 50 controlled evaluations of intelligent tutoring systems.

### Intelligent tutoring systems and learning outcomes: A meta-analysis

Ma et al., 2014. Journal of Educational Psychology. A meta-analysis of 107 effect sizes comparing learning with intelligent tutoring systems to other kinds of instruction.

## Frequently asked questions

### Does AI tutoring work?

Early randomized trials say it can when the AI is designed to teach: Kestin et al. (2025) at Harvard, Henkel et al. (2024) in Ghana and De Simone et al. (2025) in Nigeria all found gains. A plain chatbot used for answers lowered exam grades in Bastani et al. (2025).

### Has DuckyHelper been tested in a study?

Not yet. DuckyHelper is built on the learning science on this site and is trained to teach, but it has no study results of its own, and none of the numbers on this page are DuckyHelper's.

### What makes an AI tutor different from a chatbot?

A tutor is built to make you do the thinking: it asks what you think, gives hints before answers, checks your steps and brings ideas back later. A chatbot is built to answer. The research suggests the difference matters for learning.

## Sources

1. [Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15(1), 17458.](https://doi.org/10.1038/s41598-025-97652-6) (accessed 2026-10-09)
2. [Henkel, O., Horne-Robinson, H., Kozhakhmetova, N., & Lee, A. (2024). Effective and Scalable Math Support: Experimental Evidence on the Impact of an AI-Math Tutor in Ghana. AIED.](https://doi.org/10.1007/978-3-031-64315-6_34) (accessed 2026-10-09)
3. [De Simone, M., Tiberti, F., Rodriguez, M. B., Manolio, F., Mosuro, W., & Dikoru, E. J. (2025). From Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in Nigeria. World Bank Policy Research Working Paper.](https://doi.org/10.1596/1813-9450-11125) (accessed 2026-10-09)
4. [Wang, R. E., Ribeiro, A. T., Robinson, C. D., Loeb, S., & Demszky, D. (2024b). Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise. arXiv preprint (v1).](https://arxiv.org/abs/2410.03017v1) (accessed 2026-10-09)
5. [Pardos, Z. A., & Bhandari, S. (2024). ChatGPT-generated help produces learning gains equivalent to human tutor-authored help on mathematics skills. PLOS ONE, 19(5), e0304013.](https://doi.org/10.1371/journal.pone.0304013) (accessed 2026-10-09)
6. [Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122.](https://doi.org/10.1073/pnas.2422633122) (accessed 2026-10-09)
7. [Jurenka, I., et al. (2024). Towards Responsible Development of Generative AI for Education: An Evaluation-Driven Approach. arXiv (Google DeepMind tech report).](https://arxiv.org/abs/2407.12687) (accessed 2026-10-09)
8. [LearnLM Team, Google (2024). LearnLM: Improving Gemini for Learning. arXiv (Google tech report).](https://arxiv.org/abs/2412.16429) (accessed 2026-10-09)
9. [LearnLM Team, Google (2025). Evaluating Gemini in an arena for learning. arXiv (Google tech report).](https://arxiv.org/abs/2505.24477) (accessed 2026-10-09)
10. [Maurya, K. K., Srivatsa, K. A., Petukhova, K., & Kochmar, E. (2025). Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors. NAACL 2025.](https://doi.org/10.18653/v1/2025.naacl-long.57) (accessed 2026-10-09)
11. [Macina, J., Daheim, N., Hakimi, I., Kapur, M., Gurevych, I., & Sachan, M. (2025). MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors. EMNLP.](https://doi.org/10.18653/v1/2025.emnlp-main.11) (accessed 2026-10-09)
12. [VanLehn, K. (2011). The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems. Educational Psychologist, 46(4), 197-221.](https://doi.org/10.1080/00461520.2011.611369) (accessed 2026-10-09)
13. [Kulik, J. A., & Fletcher, J. D. (2016). Effectiveness of Intelligent Tutoring Systems: A Meta-Analytic Review. Review of Educational Research, 86(1), 42-78.](https://doi.org/10.3102/0034654315581420) (accessed 2026-10-09)
14. [Ma, W., Adesope, O. O., Nesbit, J. C., & Liu, Q. (2014). Intelligent tutoring systems and learning outcomes: A meta-analysis. Journal of Educational Psychology, 106(4), 901-918.](https://doi.org/10.1037/a0037123) (accessed 2026-10-09)

## Related principles

### Scaffolding: just enough help, one step at a time

What scaffolding means in learning science (Wood, Bruner and Ross; Vygotsky), why help should fade, and how DuckyHelper's hints and tutorials do it.

### Mix it up: interleaved practice

Interleaved practice in plain words: why mixing problem types helps learners tell them apart, what Rohrer and Taylor found, and what DuckyHelper does today.

### Teach one on one: the 2 sigma problem

Bloom's 2 sigma problem in plain words: what one-on-one tutoring did in the studies, what later reviews found, and how DuckyHelper tutors.

## Keep reading

### Guide, don't hand over answers

A 2025 PNAS study found that unguarded GPT-4 help raised practice grades but lowered exam grades. What it means, and how DuckyHelper guides instead.

### A model built to teach

DuckyHelper's tutor is trained on 1,000,000+ tutoring examples and inspired by 100+ learning-science papers: a tutor, not an answer machine.

### Hints before answers: guardrails for an AI tutor

White paper: what the research says about AI help that gives answers, and the guardrails DuckyHelper's tutor uses to guide learners instead.

### The learning-science library: 100+ papers

100+ real learning-science papers behind DuckyHelper, sorted by principle, each with a citation, a link to the publisher and a one-line summary.

### Why a tutor that shows you the steps works better than an app that hands you answers

Answer apps can raise practice grades and lower test scores. DuckyHelper is built the other way: it shows you the steps, gives hints first and checks your try.

## Ready to learn

DuckyHelper is ready on your Mac. Learn with Eddy on its blackboard, or with Ducky right on your screen.

[Download for Mac](https://duckyhelper.com/download/)

[See what it does](https://duckyhelper.com/features/)

The Mac app needs Apple silicon and macOS 14 Sonoma or later. There is no Chromebook, Windows or web version right now.
