A Polymarket account's X post from September 8, 2026 — 154,000+ views — claims: "Students who use AI heavily for writing assignments scored 28 points lower on science tests than peers who rarely or never use them, new OECD report warns." That's a specific, checkable number attributed to a real institution. Replies to the post immediately raised the obvious questions: is this a grading problem, not a learning problem? Isn't correlation doing a lot of work here? And doesn't "AI for writing" causing a "science" score drop sound like a garbled cross-domain claim?
We went and found the actual report. The 28-point number is real — it comes from OECD's PISA 2025 Results (Volume I), released Tuesday, September 8, 2026. But the viral post strips out nearly everything that makes the finding useful: a weekly-use sweet spot, a critical-evaluation effect that claws back nearly half the gap, and OECD's own explicit statement that the finding is an association, not proof of causation. This is worth the same rigor explainx.ai applied to Claude's Navier-Stokes rumor and the Anthropic emotion-vectors thread — check the primary source before repeating the headline number.
TL;DR
| Question | Verdict |
|---|---|
| Is the 28-point science-score gap real? | Yes — 509 (never/rarely use AI to draft writing) vs. 481 (daily users), adjusted for socioeconomic status |
| Is this a real OECD report? | Yes — PISA 2025 Results (Volume I), released September 8, 2026, 760,000+ students, 91 countries |
| Does it prove AI use causes the gap? | No — the report's own framing describes an association, not proof of causation |
| Does it control for students' prior achievement level? | Not confirmed anywhere we found — only socioeconomic status is confirmed as a control, leaving self-selection unaddressed |
| Is "more AI = worse" the actual pattern? | No — students using AI 1-2x/week outperformed both daily users and never-users |
| Does the shortcut-vs-tutor distinction matter? | Yes, explicitly — critical evaluation of AI output adds 13 points; "AI to help me learn" is the highest-scoring framing |
| Is "writing causing science scores" a mix-up? | No — it's a correlation between general AI-use habits and tested-subject performance, accurately reported, just easy to misread as causal |
What the report actually is
This is not a rumor or a misattributed working paper — it's OECD's flagship education research product. PISA (the Programme for International Student Assessment) tests 15-year-olds in science, mathematics, and reading roughly every three years, and PISA 2025 Results (Volume I) is based on a sample of over 760,000 students across 91 countries and economies, representing roughly 33 million 15-year-olds. It's the first PISA cycle assessed since generative AI reached mainstream classroom use — which is exactly why this cycle, for the first time, asked students detailed questions about how they use AI for schoolwork.
The headline number: students who said they never or almost never use AI to draft text for writing assignments scored 509 on the PISA science test. Students who use AI for that purpose daily or almost every day scored 481 — a 28-point gap, adjusted for socioeconomic status, that OECD frames as roughly a year and a half of instruction (OECD's own rule of thumb, per its coverage, is that 20 points on the PISA scale is about one year of schooling). A similar, slightly larger gap — closer to 30 points — shows up for AI use in summarizing assigned reading, and a comparable pattern appears for using AI in preliminary research on a new topic.
Context that matters for reading that number: PISA scores are scaled so the OECD average sits around 500 with a standard deviation of roughly 100. A 28-point gap is a real, non-trivial effect — about 0.28 standard deviations — but it is nowhere near the scale of, say, the entire cross-country performance spread PISA reports each cycle. It's also worth noting this cycle's broader context: average scores across all 91 economies fell to the lowest levels in PISA's history, with reading down 14 points since 2022 and math declining sharply since 2018 — a decline OECD attributes to a mix of causes including teacher shortages and shrinking attention spans, not AI use alone. Pinning that entire multi-year decline on chatbots would be its own overreach.
The nuance the viral post drops entirely
Here's where the Polymarket post's paraphrase becomes actively misleading rather than merely simplified. The report's own data does not describe a straight line where more AI use means worse scores — it describes something closer to an inverted U, with the specific mechanism of how AI gets used mattering more than whether it's used at all.
- Weekly use beats both extremes. Students who used AI about once or twice a week outperformed both students who never/rarely used it and students who used it daily. A pure "AI use is bad" reading of the 28-point number can't explain this — if frequency alone drove the effect, weekly users would land somewhere between the two extremes, not above both.
- "To help me learn" is the highest-scoring framing. When students described their AI use specifically as learning support rather than task completion, weekly users in that category scored around 500 — among the highest science scores reported for any AI-use group in the entire dataset.
- Critical evaluation adds 13 points. Among daily AI users specifically — the group with the lowest raw scores — those whose classes regularly asked them to assess AI-generated information scored 13 points higher in science than daily users who weren't taught to do that. OECD frames that as equivalent to over half a year of instruction, recovered inside the worst-performing usage tier just by teaching critical evaluation of AI output.
Put together, this is not evidence that AI use itself degrades learning. It's evidence that unreflective, task-completion AI use correlates with lower scores, while purposeful, critically-engaged AI use correlates with some of the highest scores in the same dataset — a distinction the original tweet collapses into one flat number.
Correlation vs. causation: does the concern hold up?
One of the sharpest replies under the original post argued: "correlation doing a lot of heavy lifting there, the kids who lean on it hardest were probably not acing science tests before either." That's a textbook reverse-causation and self-selection hypothesis — lower-performing students gravitating toward heavier AI use, rather than AI use causing the performance gap. Does it hold up?
We can't fully confirm or rule it out, and that itself is the honest answer. Coverage of the report's own framing states plainly that the AI-use findings show "an association rather than proof that AI use raises or lowers achievement" — language that comes from reporting on OECD's own report, not from us or from skeptical commentary. That's OECD itself declining to claim causation. What we could not confirm in any source we reviewed is whether the analysis controls for students' prior academic achievement — only socioeconomic status is confirmed as an adjustment variable. Controlling for SES rules out one confound (wealthier students having both less AI-reliance and more resources) but does nothing to rule out the specific mechanism the reply raised: a student who was already struggling in science reaching for AI more often because they're struggling, not the other way around.
That gap in what's controlled for means the self-selection hypothesis remains a live, plausible explanation for at least part of the 28-point gap — not a debunked conspiracy theory, and not a confirmed mechanism either. It's exactly the kind of caveat this outlet flagged when Anthropic's own paper explicitly declined to claim causation about what its emotion-vector data actually established, and it's worth holding OECD's report to the same standard rather than reading past the hedge because the source is prestigious.
The "grading is the problem" theory doesn't fit this specific number
A different reply suggested the real issue might be "whoever is scoring them," not the students. That's a fair general concern about AI-in-education research — a lot of studies rely on teachers grading AI-assisted homework, and grading consistency is a real confound in that context (it's part of why explainx.ai's CEPR homework study coverage leans on exam scores, not homework grades, as the more trustworthy signal). But it doesn't apply to this specific 28-point figure: PISA science scores come from a standardized, externally proctored international assessment taken without AI access, not from classroom teachers grading AI-assisted writing assignments. The measurement instrument here isn't the thing being questioned by that theory.
Is "AI for writing" causing a "science" score really a mix-up?
Also worth verifying directly, since a cross-domain causal claim (writing habits causing science outcomes) would be an odd thing for a rigorous report to assert. It isn't a mix-up, but it is easy to misread. PISA asks students how they typically use AI across several school tasks — drafting writing, summarizing reading, preliminary research, general study help — and then correlates those self-reported habits against performance on the subjects tested that cycle, science included. That's a correlation between a general AI-use pattern and standardized test performance in a different domain, not a claim that a specific writing assignment mechanically produced a specific science-test outcome. Framed that way, it's a narrower and more defensible finding than "using AI for writing homework directly damages your science knowledge" — closer to "how a student generally relates to AI as a tool tracks with broader academic performance," which is a real and interesting result, just a softer one than the viral phrasing implies.
The same shortcut-vs-tutor split shows up in prior coverage
This isn't the first study to land on the exact distinction the OECD data implies. explainx.ai covered a CEPR working paper tracking 26,811 Chinese secondary students for 30 months, which found AI use raised homework scores 18% while cutting unassisted exam scores roughly 20% — and crucially, found that roughly 80% of the exam-score decline was concentrated in students whose usage pattern matched outsourcing (very short homework time, unusually high homework scores), not in students who used AI to check their reasoning after attempting the work themselves.
The OECD data's own internal split — daily/task-completion use scoring lowest, "help me learn" framing scoring among the highest, critical-evaluation training clawing back 13 points — is the same finding from a different angle: purpose of use, not presence of use, is the variable that predicts the outcome. It also lines up with Dartmouth's Phosphor study, where a platform that graded written answers against rubrics (forcing engagement) rather than just supplying them was linked to a 0.71-1.30 standard deviation exam gain, and with the broader research cluster explainx.ai mapped in 20 studies on the "doer effect": AI that makes students produce and defend an answer tends to help; AI that hands over a finished answer tends to hurt. Three independent datasets — Chinese exam records, a Dartmouth statistics course, and now a 91-country PISA cycle — are converging on the same mechanism.
It's also consistent with Estonia's AI Leap program, a national rollout built specifically around Socratic AI tutoring instead of answer-giving, on the theory that the design of the AI interaction — not the presence of AI in the classroom — determines whether it strengthens or erodes student thinking. That's the same bet the OECD data's 13-point critical-evaluation finding supports directly.
What this means if you're a parent, educator, or student
The honest verdict is not "OECD proved AI makes students dumber" — that overclaims what even OECD itself is willing to claim about its own data. It's also not "nothing to see here, it's just correlation" — the association is real, large enough to matter (0.28 SD), reported by a credible primary source, and directionally consistent with other, more causally rigorous studies on the same underlying mechanism.
The narrower, better-supported takeaway: how AI gets used for schoolwork matters far more than whether it gets used at all. Concretely, from the report's own data plus the corroborating CEPR and Dartmouth studies:
- Daily, task-completion AI use (drafting, summarizing, research-on-demand) is the lowest-scoring pattern — treat heavy reliance on AI to produce finished writing as a warning sign, not a study habit.
- Weekly, purposeful use outperforms both extremes — "I use AI a couple times a week, specifically to help me understand something" is the profile associated with the strongest outcomes in this dataset.
- Teaching critical evaluation of AI output is a measurable, trainable skill that recovers real ground — the 13-point recovery among daily users whose classes taught them to assess AI-generated content is the single most actionable finding in the whole report, and it's a curriculum choice schools can implement without banning anything.
- Policy responses that ban AI outright — like the reported Los Angeles K-12 ban — sidestep the finding that gets students the best outcomes in this same data: not zero AI use, but weekly, critically-engaged, learning-focused use. A blanket ban forfeits the upside this report documents alongside the downside.
- Treat "28 points lower" as real but partial. It's a genuine, OECD-confirmed correlational finding worth taking seriously — not a confirmed causal indictment of AI in classrooms, and not something to wave away because a reply thread raised methodology questions. Both things are true at once, same as with the CEPR homework study and the Anthropic emotion-vector claims this outlet has fact-checked the same way.
What people are asking
Did the tweet lie? No — the 28-point figure and the writing-assignments framing are accurate to the report. What it did was strip out every piece of context that changes how the number should be read: the weekly-use sweet spot, the 13-point critical-evaluation effect, and OECD's own "association, not proof" framing. That's the difference between misquoting a source and quoting it selectively — the second is subtler and arguably more common in how viral claims spread.
Should this change how schools think about AI policy? The data argues against both extremes tested in practice right now — unrestricted AI access with no guidance, and blanket bans like LA's reported K-12 policy. It argues for the middle path Estonia's AI Leap and Dartmouth's Phosphor platform both bet on: structured, critically-engaged AI use embedded in how the tool is taught, not just whether it's allowed.
Where can I check the primary numbers myself? PISA 2025 Results (Volume I) on oecd.org is the primary source. We were not able to independently pull OECD's exact methodology section on whether prior achievement is a control variable — that specific gap is flagged above and worth verifying directly in the full report if you're citing this for further work.
Related reading on explainx.ai
- The Generative AI Learning Penalty: Homework Up 18%, Exams Down 20% — the closest comparison: a 26,811-student CEPR study finding the same shortcut-vs-tutor split
- Dartmouth's Phosphor Study: What an AI Tutor With a 0.71-1.30 SD Effect Actually Did — guardrailed AI tutoring producing the opposite result
- The Research Behind AI-Graded Quizzes: 20 Studies on Interactive Textbooks — the doer-effect research explaining why active AI use beats passive AI use
- Estonia's AI Leap: Teach Students to Think With AI, Not Instead of It — a national rollout built around the same critical-engagement principle
- Los Angeles Reportedly Bans Student AI Use Across All of K-12 — the policy-response context this report's nuance complicates
- Did Claude Solve Navier-Stokes? The Millennium Prize Rumor, Fact-Checked — same fact-check standard applied to a different viral AI claim
- Anthropic's "171 Emotion Vectors" in Claude: Fact-Checked — another primary-source check on a viral AI claim, including its own correlation-vs-causation caveat
Sources: PISA 2025 Results (Volume I) — OECD · PISA program overview — OECD · Reporting on the report's findings via Bloomberg, The Star, TNGlobal/technode.global, and Briefs.co, September 8-9, 2026
This post treats the viral X post's "28 points lower" claim as accurate to its primary source (OECD's PISA 2025 Results, Volume I) but incomplete in context, based on the report's own reported framing as of September 9, 2026. We could not independently confirm from the report itself whether the analysis controls for students' prior achievement level — that detail is flagged as unverified rather than assumed, and this post will be updated if OECD's full methodology section clarifies it.
