Homework scores went up 18%. Exam scores went down 20%. Same students, same six months, same AI tool. A working paper tracking 26,811 secondary students in China for 30 months just put a number on something teachers have suspected since ChatGPT showed up in every classroom: using AI to finish homework faster is not the same thing as learning the material — and the gap between the two shows up exactly where it matters most, on a test with the AI turned off.
This isn't the first study to find that pattern — a smaller 2024 RCT with GPT-4 tutors in Turkey found something structurally identical. But at 26,811 students and 30 months, this is by far the largest and longest-running evidence yet, and it's specific enough — down to which subjects, which students, and roughly what share of the effect is pure outsourcing — to actually change how you'd design an AI study habit, not just whether to worry.
TL;DR
| Question | Direct answer |
|---|---|
| What raised, what fell? | Homework scores +18%, homework completion time -30%, monthly exam scores -20% within 6 months, college entrance exam scores -18 to -24% |
| How many students, how long? | 26,811 students, grades 7-12, tracked for 30 months across a county in central China |
| Who ran the study? | David Strömberg, Victor Lei, and Yanhui Wu — Stockholm University and the University of Hong Kong — published via CEPR (DP21577) |
| What actually caused the exam drop? | ~80% of the loss is concentrated in students whose usage pattern matches "outsourcing" — very short homework time plus high homework scores |
| Does this mean all AI tutoring is bad? | No — guardrailed tools (hint-only prompts, answer grading instead of answer-giving) show the opposite effect in other studies |
| What should I actually do? | Use AI to check work and explain concepts after attempting the problem yourself, never to produce the answer you hand in |
What the study actually measured
The paper — "The Generative AI Learning Penalty: Evidence from Chinese Secondary Education", CEPR Discussion Paper DP21577 — used 30 months of panel data on 26,811 students in grades 7 through 12 across a county in central China, tracking both homework performance and exam performance as generative AI tools got adopted at different rates across classrooms. That panel structure is what makes the finding stronger than a single before/after comparison: the researchers could watch the same students' homework and exam trajectories diverge over time as AI use increased, rather than just comparing two snapshots.
The headline numbers:
- Homework scores rose 18%, and completion time fell 30% — AI made homework faster and higher-scoring, exactly as you'd expect.
- Monthly exam scores fell 20% within six months of AI adoption — a test taken without AI access, on the same material.
- College entrance exam scores fell 18-24%, with the full size of the penalty only showing up after about two years — the highest-stakes exam in the dataset took the longest to reveal the damage.
- Losses were largest in social science subjects, then STEM, then languages, and were especially large for junior-grade students, high-achieving students, and boys.
The mechanism: outsourcing, not AI itself
The most important finding isn't the headline gap — it's what's driving it. The researchers found that roughly 80% of the exam-score decline was concentrated in students whose behavior pattern matched homework outsourcing: exceptionally short homework completion time combined with unusually high homework scores. That combination is the signature of a student pasting a question into a chatbot and copying the answer back, rather than attempting the problem and using AI to check or explain it.
That distinction matters because it means the "penalty" isn't a property of generative AI as a technology — it's a property of how it gets used. A student who spends the same amount of time on a problem, but uses AI to get unstuck or verify their reasoning, doesn't show the same signature in the data as one who outsources the whole task. The researchers' own framing, echoed across the coverage: for students, completing homework efficiently was never the goal — learning from it was, and generative AI made it trivially easy to optimize for the wrong variable.
This is the same mechanism explainx.ai has already covered on the other side of the age range: our AI-driven de-skilling piece covers an Anthropic RCT finding a 17% comprehension deficit in professional developers who leaned on AI coding assistants instead of writing code themselves. Same shape of result, same underlying cause, different population — practice that looks productive in the moment quietly erodes the skill it was supposed to build.
Guardrails change the outcome — this isn't the first study to show it
This isn't a lone data point, and it isn't proof that AI-assisted learning is doomed. A 2024 randomized controlled trial with nearly 1,000 Turkish high school math students — run by Hamsa Bastani, Osbert Bastani, and colleagues, and later published in PNAS — split students into three groups: unrestricted GPT-4 access, a "GPT Tutor" version that gave hints instead of direct answers, and a no-AI control. The result was almost a preview of the Chinese study's finding, compressed into one semester: unrestricted GPT-4 access boosted in-session practice scores 48%, but once access was removed, those same students scored 17% worse than the control group on an unassisted test. The hint-only GPT Tutor group, given access to the identical underlying model, showed the improvement without the same collapse when AI was removed — the guardrail, not the model, was the deciding factor.
explainx.ai's own coverage of Dartmouth's Phosphor study points the same direction from a different angle: a platform that graded written answers against rubrics instead of supplying them was linked to a 0.71-1.30 standard deviation final exam gain, with 90.2% voluntary adoption — while the platform's actual open-ended chatbot barely got used at all. The lesson across all three studies is consistent: AI that grades, hints, or quizzes tends to help; AI that answers tends to hurt, and the difference is entirely in the guardrails, not the underlying model capability.
What this means if you're using AI to study
The practical takeaway isn't "stop using AI for homework" — it's stop using it as an answer machine and start using it as a tutor. explainx.ai's family framework for AI and homework puts this as a single line simple enough for a ten-year-old to apply: AI can teach you. AI cannot be you. A tutor explains a concept differently, checks your work, and quizzes you. A ghostwriter produces the assignment while you watch. The Chinese study's 80%-outsourcing finding is empirical backing for exactly that distinction — the students who used AI like a ghostwriter are the ones who lost ground; the ones who didn't largely didn't.
A concrete self-check that follows from the guardrail research above: the explain-it-back test. If you can't explain your own answer with the AI tool closed, you haven't learned the material — you've borrowed a correct-looking output that will not be there for you on the exam. That's precisely the mechanism the study's "outsourcing signature" (short time, high score) is detecting at scale.
This is also the design principle behind Melo, explainx.ai's own AI learning copilot: its Quiz, Practice, and Explain Back modes are built to make you produce the answer and defend it, the same shape as Phosphor's constructed-response questions and the Turkish study's hint-only GPT Tutor, rather than a chat window that hands you a finished answer to paste. If your current AI habit is "ask, copy, submit," swapping in a Quiz or Explain Back pass on the same material is a low-effort way to get the homework speed-up without the exam-score penalty this study documents. Our interactive learning pathways apply the same principle at the course level — structured practice with checks, not just a chat box.
What people are asking
"Is this specific to China's education system?" The mechanism isn't — homework-outsourcing behavior and the practice/test gap it creates aren't unique to any one country's curriculum. What's specific to this study is its scale and duration (26,811 students, 30 months), which is why it's useful as the largest data point we have, not the only one — it lines up with the smaller Turkey RCT and with Dartmouth's guardrailed-tool result.
"Should schools just ban AI for homework then?" The study doesn't test that policy directly, and the guardrail research above suggests banning AI outright forfeits the 18% homework-score and 30% time-savings benefit for no reason — a hint-only or answer-checking tool captured the upside in the Turkey RCT without the downside. The more defensible policy, backed by this data, is restricting what kind of AI access students get (hints and checking vs. direct answers), not whether they get any at all.
"How long before the damage shows up?" Fast on lower-stakes assessments — monthly exam scores dropped 20% within six months — but the full size of the effect on the highest-stakes exam (college entrance) only appeared after about two years. That lag is itself a warning: a student (or a school) could look at six months of rising homework scores and falling AI dependency-awareness and conclude things are fine, well before the real cost shows up on the exam that counts most.
Related on explainx.ai
- Dartmouth's Phosphor study: what a guardrailed AI tutor actually did
- AI and homework: house rules that actually work
- AI-driven de-skilling: the developer version of the same problem
- Introducing Melo: explainx.ai's AI learning copilot
- Melo's generative UI for real-time learning
- Introducing interactive AI learning pathways
- ChatGPT for teens: safety features and Study Mode
Official sources: CEPR Discussion Paper DP21577 · SSRN preprint · Bastani et al., "Generative AI Without Guardrails Can Harm Learning," PNAS
Figures in this post reflect the CEPR working paper and PNAS study as published; working papers can be revised before final peer-reviewed publication, and effect sizes from observational and quasi-experimental data should be read as strong evidence, not absolute certainty.
