Guide, don't hand over answers
In a 2025 field experiment in a high school, students who practiced math with a plain GPT-4 chat did better during practice, then 17% worse on the exam once the AI was taken away, compared with students who never had it. A version built to give hints instead of answers mostly avoided the harm. DuckyHelper is built the second way: it guides you to the answer.


Hints before answers
In the film, a learner asks for the answer and the PNAS page shows what that cost. DuckyHelper is built the other way:
- Try first, then guide (Mac and web)
- Hints, not answers (Mac app)
- A nudge on a wrong answer (Mac and web)
- Steps you have reached, not the whole solution (Mac and web)
- Questions of fact still get answers (Mac and web)
What the research found
Hamsa Bastani and colleagues ran a field experiment with high school math students. Some practiced with "GPT Base", a plain ChatGPT-style interface; some with "GPT Tutor", the same model with prompts designed to protect learning; and some with no AI. During practice, both AI groups did better: grades rose 48% with GPT Base and 127% with GPT Tutor.
Then the AI was taken away for the exam. Students who had used GPT Base did worse than students who never had it, a 17% reduction in grades. In the authors' words, “unfettered access to GPT-4 can harm educational outcomes.”
The guarded version mostly avoided this. In the authors' words:
These negative learning effects are largely mitigated by the safeguards in GPT Tutor
(Bastani et al., 2025)
The authors' explanation is that without guardrails, students used the AI as a crutch during practice and then could not do the work alone. Other research points the same way. Language models solve math well but tend to give the solution away too early when they tutor:
While models like GPT-3 are good problem solvers, they fail at tutoring
(Macina et al., 2023a)
And long before chatbots, students who used a tutor's hints just to get to the answers ("gaming the system") learned much less:
students who frequently game the system score substantially lower on a post-test than students who never game the system
(Baker et al., 2004)
Feature by feature
Try first, then guide
On homework and anything graded, it does not just hand you the final answer. You take a first try, then it guides you one step at a time. Really stuck? It works through a similar problem with you, and then you do yours. Mac and web.
Hints, not answers
Stuck on a problem in the Mac app? Press Hint. A ladder opens one rung at a time: a nudge, a bigger hint, the worked step, and the answer last. It holds the answer back until you ask for it a second time. Mac app.
A nudge on a wrong answer
In a quiz, a wrong answer gets a one-sentence nudge, and a hint request gets a hint without the answer. Mac and web.
Steps you have reached, not the whole solution
On homework, worked steps stop where you are and add a hint for your next step, instead of the whole solution. On Eddy's blackboard, math is worked out step by step, not just announced. Mac and web.
Questions of fact still get answers
Guiding is for problems you are learning to solve. Ask a simple fact and you get a clear, short answer, plus one small thing to try next. Mac and web.
What this page does not claim
These studies tested other people, other schools and other tools. DuckyHelper has not been tested in a study like these, so none of these results are DuckyHelper's results. They are the reasons behind how it teaches.
The papers on this page
Frequently asked questions
Does using ChatGPT for homework hurt learning?
It can. In Bastani et al. (2025), high school students who practiced with a plain GPT-4 chat did better during practice but 17% worse on the exam once it was taken away, compared with students who never had it. A version with learning safeguards largely avoided that harm.
Will DuckyHelper ever just tell me the answer?
For a plain question of fact, yes, briefly. For homework and practice problems it guides you first: it asks what you think, gives hints one at a time and shows similar worked examples. In the Mac app's hint ladder the answer is the last rung.
Is a tutor that withholds answers slower?
Sometimes, on purpose. The research above suggests that the effort of working it out is part of how the learning happens. DuckyHelper keeps that effort small with quick hints and steps instead of long waits.
Sources
- Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122. (accessed 2026-10-09)
- Macina, J., Daheim, N., Chowdhury, S., Sinha, T., Kapur, M., Gurevych, I., & Sachan, M. (2023a). MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning Problems. Findings of EMNLP 2023. (accessed 2026-10-09)
- Baker, R. S., Corbett, A. T., Koedinger, K. R., & Wagner, A. Z. (2004). Off-Task Behavior in the Cognitive Tutor Classroom: When Students "Game the System". CHI. (accessed 2026-10-09)
- Sonkar, S., Ni, K., Chaudhary, S., & Baraniuk, R. (2024). Pedagogical Alignment of Large Language Models. Findings of EMNLP. (accessed 2026-10-09)
- Dinucu-Jianu, D., Macina, J., Daheim, N., Hakimi, I., Gurevych, I., & Sachan, M. (2025). From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement Learning. EMNLP. (accessed 2026-10-09)
- LearnLM Team, Google (2024). LearnLM: Improving Gemini for Learning. arXiv (Google tech report). (accessed 2026-10-09)
- Wang, R. E., Ribeiro, A. T., Robinson, C. D., Loeb, S., & Demszky, D. (2024b). Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise. arXiv preprint (v1). (accessed 2026-10-09)
Related principles

Ready to learn
DuckyHelper is ready on your Mac. Learn with Eddy on its blackboard, or with Ducky right on your screen.