Insights

Does Studying With AI Help? Practice Up 48%, Exam Down 17%

7 min read#ai#learning#study-habits#productivity

Who this is forReaders wondering whether studying with AI actually improves skills, or how much AI they should allow their kids or team members to use.

TL;DR: Studying with AI can raise your scores right away, but on an exam taken without AI, you may score lower than students who studied without it. This result comes from an experiment with about 1,000 high school students. However, one group avoided the decline, and the difference lay not in the tool but in how it was used. This post walks through that experiment and the principles I put into practice to keep AI from doing my thinking.

Learning something has become remarkably easy. When I don’t know something, I ask an AI and get an answer immediately. When I open YouTube, the algorithm keeps serving up videos that seem worth learning from. But then a strange thing happens: I feel like I learned something, yet when I try to do it on my own, nothing stays with me.

I wondered if it was just me, but there was an experiment that measured this phenomenon precisely.

Contents

  1. The GPT-4 Experiment With About 1,000 High Schoolers
  2. The Problem Is Delegation, Not AI
  3. So I Deleted the Recommendation Feed Apps
  4. Want to Read Further?

The GPT-4 Experiment With About 1,000 High Schoolers

Researchers at the University of Pennsylvania ran an experiment with about 1,000 students at a high school in Turkey. During math practice sessions, they divided the students into three groups. One group used a standard chatbot-style GPT-4. Another used a tutor-style GPT-4 designed with guardrails so that it would not give answers directly. The rest studied without AI. Later, all students took the exam without AI.

Experiment results: everyone's practice scores rose, but the exams split

  1. Standard chatbot-style GPT-4 Practice +48% / Exam -17% Students who took answers directly scored lower on the exam than students who studied without AI once the AI was gone
  2. Guardrailed tutor-style GPT-4 Practice +127% / Exam: no difference The group given only hints and made to solve the answers themselves had the largest practice gain, and on the exam it was not statistically distinguishable from the group that studied without AI
  3. Studying without AI Baseline The benchmark against which the exam scores of the two AI groups were compared

What made the difference was not whether AI was used, but who did the thinking

Here is what the numbers look like as bar lengths.

Same GPT-4, but the guardrails decided the outcome

Difference in score compared with the group that studied without AI, %

Practice problems (time spent solving with AI) Standard chatbot-style +48 Guardrailed tutor-style +127 Exam (results taken without AI) Standard chatbot-style -17 Guardrailed tutor-style 0, no statistically significant difference from the control group 0 +50 +100

Source: Bastani et al., PNAS vol. 122, no. 26 (2025), Table 1

Because the original chart's license (CC BY-NC-ND) does not allow copying, I redrew it using the figures from the paper. Original: PNAS original (free full text)

One caveat: even the guardrailed tutor-style group did not raise exam scores. It only eliminated the decline. In the end, the ability that shows up in the exam room only builds up in proportion to the thinking your own head does during practice.

The Problem Is Delegation, Not AI

This experiment is not an argument against studying with AI. The same tool hurt one group and didn’t hurt the other. The dividing line is whether you hand over the entire thought process, or get help only at the point where you’re stuck.

Of course, this study comes from one high school in Turkey and one subject, math, so it can’t be generalized as is. Still, I think this distinction also applies to how we consume AI and content. Letting a recommendation algorithm decide even what to watch, or pasting in an answer you received whole, both hand over the role of the thinker. They share the same structure.

Delegating thinking Growing thinking
Give the AI the whole problem and copy down the answer Try it first, then ask only about the point where you’re stuck
Leave what to watch to the recommendation feed Decide what to learn first, then find it yourself
Read only the AI summary and stop Use summaries to choose, then read the original yourself
Use whatever result the AI produces as is Rewrite it in your own words and find where it doesn’t fit

The left side of the table is more comfortable. And while you stay on the left, the feeling of “I learned something” keeps coming. The most chilling passage in the paper is here. The researchers wrote that the students who copied down answers did not notice on their own that this method was harming their learning. A feeling does not guarantee ability.

So I Deleted the Recommendation Feed Apps

Once I set that principle, what I did next was not dramatic. I deleted the apps I had been using to consume recommendation feeds. To be honest about the trigger, it came after I spent several days in bed with a bout of gastroenteritis, soaked in YouTube Shorts. Interesting videos kept coming endlessly, and when I looked back a few days later, there was nothing left from that time.

Time spent watching what the recommendations chose

  1. Who picks Algorithm The goal is time spent on the platform, and your growth is not the goal
  2. What's left afterward Feeling of having learned If you try it on your own, nothing is left

Time spent making what I chose

  1. Who picks Me Decide what to learn first, and use tools only in that direction
  2. What's left afterward Work and skill Tangible results remain, like an essay or an automation you built

Johann Hari, in Stolen Focus, points to an industry structure in which all kinds of techniques are deployed to capture people’s attention. Recommendation feeds are designed to get you to swipe one more time, and time spent becomes ad revenue. Inside that structure, continuously receiving “videos worth learning from” is closer to consumption wearing the face of study.

If you want to move from the place where you learn to the place where you make, I recommend starting with small projects that leave something tangible behind, such as Your First Automation with n8n on this blog. If the unfamiliar terms of AI automation feel like a wall, there is a glossary series that starts with What Is an API?.

Want to Read Further?

  • Generative AI without guardrails can harm learning (PNAS, 2025) The original paper for the experiment introduced in the body. The link goes to the PMC version, which has free full text.
  • Johann Hari, Stolen Focus. A book on the industry structure that captures attention. It reads better after you’ve deleted your recommendation feeds.

I checked the experiment figures in this post directly against the original paper, published in PNAS in June 2025 (vol. 122, no. 26; researchers at the University of Pennsylvania; high school students in Turkey), on August 31, 2026. The standard chatbot-style group showed +48% on practice and scored 17% lower than the control group on the exam. The guardrailed tutor-style group showed +127% on practice and no statistically significant difference from the control group on the exam.

Frequently asked questions

Does studying with a standard AI chatbot lower exam scores taken without AI?
In the experiment with about 1,000 high school students in Turkey, the standard GPT-4 chatbot group gained 48% on practice problems but scored 17% lower on the exam taken without AI than the group that studied without AI.
Did the guardrailed tutor-style GPT-4 raise exam scores?
No. The tutor-style GPT-4, which gave hints instead of answers, eliminated the exam decline but did not raise exam scores above the group that studied without AI.

Want the full system? The Claude Code & Codex Skills guidebook collects the skills and subagents behind this blog, from $19.