Experimental Evidence on the Learning Impact of Generative AI
Zara Contractor, Germán Reyes
econ.GN, cs.HC
2026-07-10
Middlebury RCT, n=211: GPT-4o access raised an immediate quiz 7.2 percentage points (0.28 SD). A week later, 3.9 remained. Tutor-style gains lasted; ghostwriting did not.
Students already use stock chatbots on coursework. Whether that produces knowledge is much less clear. Many experiments modify the system, adding teacher-written hints or blocking direct answers, or they compare AI with active learning and human tutoring. Others let students choose how long to work, so time on task and learning per minute move together.
This experiment freezes the clock and estimates learning productivity only. It was run in proctored labs at Middlebury College, and every scored assessment happened after AI and other sources were taken away.
Of the undergraduates recruited, 211 attended session one and 205 attended both. Each student was randomized to blockchain, carbon capture, or CRISPR gene editing. Baseline self-rated knowledge was 1.5 out of 10, and a five-item quiz was answered correctly 30.3% of the time. The learning block lasted at most 35 minutes and included an analytical essay of about 500 words plus a primer of 1,248 to 1,382 words, followed by a new five-item quiz.
About a week later (mean gap 6.98 days) everyone worked with no AI and no outside sources: a ten-item quiz and a 20-minute essay. Students had not been told those assessments were coming. Five percent looked anything up in between, and none reported studying.
The AI-allowed lab could use any generative model. A GPT-4o account was already signed in, with Wikipedia and the library open. The AI-forbidden lab was told not to use ChatGPT, Claude, or similar tools, and had Google, Wikipedia, and the library instead. Proctors, screen captures, and an exit survey checked compliance. Pay was $50 for both sessions and $5 for session one alone, plus a lottery: 30 tickets worth $100, earned from correct answers and essay points.
The main estimate is the effect of being allowed to use AI, whether or not a student opened it. Assignment also instruments actual use.
Activity in the provided ChatGPT account rose 65.9 percentage points, and self-reported use of any generative AI rose 60.9 points, from a near-zero control baseline. In the logs, 56% of allowed students asked for explanations, 29% for drafts, 23% for edits, and 16% for summaries. Among users, 88% called the tool helpful. Time in the learning block barely moved (control mean 32.6 minutes). The share spent writing fell 5.3 points from 53.4%, and the share spent reading and searching rose 4.4 points from 44.5%. Enjoyment rose 0.58 points from a control mean of 5.22 on a 0–10 scale.
| Outcome | Control | AI-allowed vs control |
| Immediate quiz | 56.3% | +7.2 pp (0.28 SD) |
| Quiz one week later | 50.0% | +3.9 pp (0.21 SD) |
| Immediate essay | 5.50/10 | +0.45 points (0.34 SD) |
| Unaided essay one week later | — | +0.33 points (0.24 SD) |
| Share of text flagged as AI | 6.6% | +22.2 pp; +0.7 pp a week later |
Students whose use was shifted by assignment scored 10.4 points higher on the immediate quiz (0.41 SD). Self-rated knowledge rose by about the same amount in both groups. A 0.28 SD effect sits above the 0.10 SD median from 747 education RCTs with standardized tests, and next to the 0.29 SD pooled effect of structured human tutoring. Across 14 other AI-and-learning experiments, the random-effects mean is 0.19 SD. Within-group essay similarity rose by 0.003 from a control mean of 0.788. On the delayed essay, use of evidence is the largest dimension, at 0.35 SD.
What lasted depends on how students used the tool, and that choice was not randomized again. Of users, 61% only asked AI to work with them (augmentation), 9% had it write the essay (automation), and 27% did both. Automation users' essays scored 1.50 SD above control in session one and 0.14 SD below control a week later. Augmentation users scored 0.34 SD higher on the immediate quiz and 0.24 SD higher a week later. Mixed users scored 0.58 SD higher immediately; the delayed 0.28 SD is imprecise.
Control students, who never received access, predicted a 25.5 point gain on their own quiz, 6.5 times the 3.9 point effect measured a week later. Students who had access predicted 3.0 points.
A college can permit AI. It cannot require every student to use it, so the permission effect is the policy parameter. Stock GPT-4o, with no teacher-written guardrails, still raised what students could do once the tool was gone. Of the immediate 7.2 point quiz gain, 3.9 points were still there a week later.
The part that lasted is tutor-style use. Ghostwritten essays looked better on the day and lost that edge once AI was unavailable. An effort model in the paper fits the split: tutoring raises the knowledge return to reading and search, while delegation removes the grade reward for writing the essay oneself, so effort on both activities falls.
The lab offered little else to do, and 35 minutes was a hard cap, so AI did not cut total study time. On a real assignment, students can spend the saved minutes somewhere else. Whether that reallocation wipes out the productivity gain is not measured. The paper says the results are not evidence that broad adoption will raise learning overall.
The use-type split is descriptive. Pure automation is 9% of users, and their delayed quiz estimate, −0.24 SD, has a p-value of 0.538. The immediate quiz is also slightly contaminated: proctor-observed rule violations rose 5.6 points and self-reported violations 9.6 points, enough to account for roughly 17% to 22% of that effect in a back-of-the-envelope calculation. Dropping violators leaves estimates between 0.15 and 0.27 SD. Nobody broke the rules in session two, but the delayed estimates are marginal: p = 0.068 for the 0.21 SD quiz effect and p = 0.083 for the 0.24 SD essay effect.
Mean GPA is 3.68 and mean SAT about 1386, campus AI adoption was already above 80%, and the sample over-represents women and first-years. Three technical topics, one week of follow-up. Headline essay scores average Prolific graders who hold a master's or PhD with an LLM grade, and those graders may not know the subject.