What the evidence says about AI tutors
Early trials are promising when the AI is built to teach: a Harvard physics tutor, a math tutor in Ghana and an English program in Nigeria all found learning gains. When students got a plain chatbot, a 2025 PNAS study found exam grades fell once it was gone. The design decides it. DuckyHelper is built to teach, and has not itself been tested in a trial yet.
“students learn significantly more in less time when using the AI tutor”


Built like the tutors that helped
In the film, a learner asks for the answer and the PNAS page shows the cost of a plain chatbot. DuckyHelper follows the designs that worked in these trials:
- Built to teach, not to answer (Mac and web)
- Trained to teach (Mac and web)
- Learning science in the tools (Mac and web)
What the research found
When the AI is built to teach
At Harvard, Gregory Kestin and colleagues compared an AI tutor, designed with the same teaching practices as the course, against an in-class active learning lesson. Students learned significantly more, in less time, with the AI tutor (the line quoted at the top of this page).
In Ghana, about 500 students used Rori, a math tutor on WhatsApp, for one hour a week during study hall. Their growth scores were substantially higher, an effect size of 0.36 (Henkel et al., 2024). In Nigeria, a six-week after-school program with Microsoft Copilot (GPT-4) raised English results:
The effect on English, the main outcome of interest, was of 0.23 standard deviations
(De Simone et al., 2025)
AI can also help human tutors. In Tutor CoPilot, students whose tutors had AI suggestions were 4 percentage points more likely to master topics, and the tutors were
less likely to give away the answer to the student
(Wang et al., 2024b)
ChatGPT-written hints produced learning gains like human-written hints in Pardos and Bhandari (2024).
When the AI just answers
unfettered access to GPT-4 can harm educational outcomes
(Bastani et al., 2025)
In that study the plain chatbot raised practice grades and cut exam grades by 17%, while the version with learning safeguards mostly avoided the harm. See Guide, don't hand over answers.
Measuring teaching, not just answering
Research groups are now testing whether models teach well, not only whether they are right. Google's LearnLM work found pedagogy-tuned models preferred by educators, and an evaluation taxonomy for AI tutors separates the two jobs:
highlighting which LLMs are good tutors and which ones are more suitable as question-answering systems
(Maurya et al., 2025)
subject expertise, indicated by solving ability, does not immediately translate to good teaching
(Macina et al., 2025)
Older research on computer tutors is the long view: reviews found intelligent tutoring systems raised test scores (Kulik and Fletcher, 2016; Ma et al., 2014), nearly as much as human tutors in VanLehn (2011).
Feature by feature
Built to teach, not to answer
A tutor, not an answer machine: it asks, hints and checks, following the safeguards that worked in the trials above. Mac and web.
Trained to teach
DuckyHelper's tutor is trained on 1,000,000+ tutoring examples and inspired by 100+ learning-science papers, the library on this site. Mac and web.
Learning science in the tools
Quizzes, spaced flashcards, worked steps, hints and drawing are each tied to a principle on this site. Mac and web.
What this page does not claim
No study of DuckyHelper exists yet. Every number on this page belongs to the cited study and the tool it tested, not to DuckyHelper.
The papers on this page
From Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in Nigeria
Frequently asked questions
Does AI tutoring work?
Early randomized trials say it can when the AI is designed to teach: Kestin et al. (2025) at Harvard, Henkel et al. (2024) in Ghana and De Simone et al. (2025) in Nigeria all found gains. A plain chatbot used for answers lowered exam grades in Bastani et al. (2025).
Has DuckyHelper been tested in a study?
Not yet. DuckyHelper is built on the learning science on this site and is trained to teach, but it has no study results of its own, and none of the numbers on this page are DuckyHelper's.
What makes an AI tutor different from a chatbot?
A tutor is built to make you do the thinking: it asks what you think, gives hints before answers, checks your steps and brings ideas back later. A chatbot is built to answer. The research suggests the difference matters for learning.
Sources
- Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15(1), 17458. (accessed 2026-10-09)
- Henkel, O., Horne-Robinson, H., Kozhakhmetova, N., & Lee, A. (2024). Effective and Scalable Math Support: Experimental Evidence on the Impact of an AI-Math Tutor in Ghana. AIED. (accessed 2026-10-09)
- De Simone, M., Tiberti, F., Rodriguez, M. B., Manolio, F., Mosuro, W., & Dikoru, E. J. (2025). From Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in Nigeria. World Bank Policy Research Working Paper. (accessed 2026-10-09)
- Wang, R. E., Ribeiro, A. T., Robinson, C. D., Loeb, S., & Demszky, D. (2024b). Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise. arXiv preprint (v1). (accessed 2026-10-09)
- Pardos, Z. A., & Bhandari, S. (2024). ChatGPT-generated help produces learning gains equivalent to human tutor-authored help on mathematics skills. PLOS ONE, 19(5), e0304013. (accessed 2026-10-09)
- Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122. (accessed 2026-10-09)
- Jurenka, I., et al. (2024). Towards Responsible Development of Generative AI for Education: An Evaluation-Driven Approach. arXiv (Google DeepMind tech report). (accessed 2026-10-09)
- LearnLM Team, Google (2024). LearnLM: Improving Gemini for Learning. arXiv (Google tech report). (accessed 2026-10-09)
- LearnLM Team, Google (2025). Evaluating Gemini in an arena for learning. arXiv (Google tech report). (accessed 2026-10-09)
- Maurya, K. K., Srivatsa, K. A., Petukhova, K., & Kochmar, E. (2025). Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors. NAACL 2025. (accessed 2026-10-09)
- Macina, J., Daheim, N., Hakimi, I., Kapur, M., Gurevych, I., & Sachan, M. (2025). MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors. EMNLP. (accessed 2026-10-09)
- VanLehn, K. (2011). The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems. Educational Psychologist, 46(4), 197-221. (accessed 2026-10-09)
- Kulik, J. A., & Fletcher, J. D. (2016). Effectiveness of Intelligent Tutoring Systems: A Meta-Analytic Review. Review of Educational Research, 86(1), 42-78. (accessed 2026-10-09)
- Ma, W., Adesope, O. O., Nesbit, J. C., & Liu, Q. (2014). Intelligent tutoring systems and learning outcomes: A meta-analysis. Journal of Educational Psychology, 106(4), 901-918. (accessed 2026-10-09)
Related principles

Ready to learn
DuckyHelper is ready on your Mac. Learn with Eddy on its blackboard, or with Ducky right on your screen.