The Homework Score Is No Longer Evidence of Learning

students learning

Two 2026 studies show that AI raises the numbers teachers watch and lowers the learning those numbers were built to represent. The measurement problem belongs to everyone building for schools.

A student who used to spend 64 minutes on an assignment now finishes it in 45. Her homework score goes up by 18 percent. Every dashboard in the school reads that as progress. The teacher sees a strong submission. The parent sees a good week. The platform reports higher completion and higher scores, and someone in a product meeting calls it engagement.

Six months later the same student sits a closed-book exam and scores 20 percent lower than she would have before.

That is the finding at the centre of a new study of 26,811 Chinese secondary students, and it should change what we count. The homework score has stopped being evidence of learning. Almost every education product and school reporting system still treats it as though it is.

What the study found

David Strömberg of Stockholm University, Victor Lei and Yanhui Wu of the University of Hong Kong tracked 30 months of administrative records from students in grades 7 to 12 in a county in central China. The data cover nine subjects and link weekly homework scores and completion time to monthly closed-book exams, county-wide exams, and the high school and college entrance examinations. The work is published as CEPR Discussion Paper 21577 and is also available on SSRN. The authors set out the argument for a general audience in a VoxEU column published on 4 September 2026.

The researchers needed to know when each student started using generative AI. They got it from surveys distributed by teachers, who asked students to check the registration dates on the AI tools they had used. By June 2025 roughly 80 percent of students reported using generative AI, up from almost none in September 2022. The tools students chose were Doubao and DeepSeek, general purpose assistants with no educational design in them.

Because students adopted AI at different times, the researchers could compare what happened to a student after adoption against students who had not yet adopted. Before adoption the two groups had almost identical scores and followed similar trends. That staggered difference-in-differences design carries far more weight than a simple comparison of AI users against non-users, because it watches the same student cross the line.

The productivity gain is real. Within six months of adoption, average homework completion time fell from 64 minutes to about 45, and homework scores rose 18 percent against the pre-adoption mean.

The learning went the other way. Monthly closed-book exam scores fell 20 percent within six months. High school entrance exam scores fell 24 percent and college entrance exam scores fell 18 percent, with the full penalty arriving only after about two years.

The gap between those two timelines matters. Monthly exams test recent material. Entrance exams integrate material built up over years. The authors draw the obvious conclusion, which is that short studies systematically understate the long-run cost. Anyone holding a six-week pilot showing that AI helps should read that sentence twice.

The finding that should worry us most

Among students who did not use AI, the study found the relationship every teacher assumes: higher homework scores went with higher exam scores. Homework was doing its job as a signal.

Among AI users, that relationship inverted. Students with higher homework scores were more likely to perform worse on exams.

Read that again as a builder rather than as a teacher. The number the system collects has not just become noisy. It has changed sign. The authors put it plainly. A rapid increase in homework scores now sends a warning signal of learning decline.

Every progress bar, every completion badge, every weekly report we ship runs on that number.

Two ways to use the same tool

The losses were not spread evenly across AI users. After six months, about four in every five AI-using students were finishing assignments in under 50 minutes, faster than the fastest students who used no AI at all, and posting homework scores that matched what the tools themselves can produce. Their exam scores were markedly worse. The pattern fits students handing substantial parts of the work to the machine.

The remaining one in five kept spending roughly as much time on homework as non-users. Their exam scores looked like the exam scores of non-users. They were still using AI, as their higher homework scores show. They were just not using it to skip the work.

The tempting policy response is to mandate longer homework time. The authors close that door themselves. Students who use AI differently also differ in motivation, in what they know about the costs, and in how skilled they are with the tools. Time on task is a symptom of the choice, not the cause of the outcome.

The damage also concentrates in places that cut against intuition. Exam scores fell about 27 percent in social science subjects such as history and politics, about 22 percent in STEM, 17 percent in English and 9 percent in Chinese. Losses were disproportionately large for junior students, for boys, and for students who were performing well before they adopted AI. Labour market research has found that AI tends to compress the skill distribution by lifting less skilled workers. In this school setting the distribution compresses for the opposite reason. The penalty lands hardest on the strongest students.

A second dataset, a different method, the same direction

The OECD released PISA 2025 Results (Volume I) this month, covering more than 760,000 students across 91 countries and economies. It is the first PISA cycle run after generative AI became widely available.

Students who used AI for specific schoolwork tasks, summarising a text, doing preliminary research, drafting a written assignment, scored lower in science than students who did not, by around 20 points on average. The OECD treats 20 points as roughly one year of schooling. The widest gap appeared in summarising assigned reading, where daily users scored close to 30 points below non-users. For drafting written assignments, students who never or almost never used AI averaged 509 against 481 for daily users, and the 28 point difference held after adjusting for socioeconomic background.

Then comes the part that stops this from being an argument against AI. Students who used AI about weekly for the general purpose of helping them learn performed similarly to non-users. Students who used it regularly for general purposes and who also had opportunities at school to build AI literacy tended to score slightly higher in science.

PISA is observational. It compares what students reported doing against how they performed. It does not show that AI use caused the difference, and the OECD says so. Students who reach for a chatbot may already differ from those who do not in motivation, in prior achievement, in the help available to them at home, and in the kind of work their schools set.

That caveat is why the two studies matter more together than either does alone. One is a causal design following the same students across adoption in one county in China. The other is an association across 91 systems. Different methods, different scales, different parts of the world, pointing the same way. Using AI to produce the work costs learning. Using AI to support the work does not appear to.

A third result closes the gap between them. In a randomised controlled trial reported in PNAS in 2025, Bastani and colleagues found that a generative AI tool supplying direct answers to homework questions improved students’ performance on practice problems but lowered their closed-book mathematics exam scores by 17 percent. A separate trial by Kestin and colleagues, published in Scientific Reports the same year, found that a guarded and pedagogically designed AI tutor improved learning. The tool is not the variable. The design of the tool is.

The protection is unevenly distributed

The OECD adds one more finding that deserves its own weight. The learning opportunities that build AI literacy, the ones that came with slightly higher science scores, are more common among socioeconomically advantaged students.

So the thing that appears to protect learning is itself concentrated where advantage already sits. The tools spread everywhere at once, at almost no cost, on any phone. The teaching that makes the tools safe spreads slowly and expensively and unevenly. Systems with the least capacity to teach students how to judge what a model gives them carry the largest exposure to the penalty.

Any product that ships an AI feature into a school without shipping the literacy alongside it is relying on a resource the school may not have.

What we should be measuring

The measurement failure is ours. As a field we have been reporting assisted task performance and calling it learning.

Higher homework scores and shorter completion times are not evidence of learning. They never were sufficient, and in the presence of a general purpose model they now point the wrong way. What counts as evidence is delayed retention, closed-book performance, transfer to problems the student has not seen, and performance without the assistant present. Those are the numbers that should appear on the dashboard, and they are harder and slower to collect, which is exactly why they have not been.

Design follows measurement. A system built to raise the numbers above has to make the student retrieve, reason, explain, generate an answer and revise it. AI can give hints, feedback, critique and adaptive practice, and it can do those things well. It must not do the target cognitive work. The design objective is better independent performance later, not faster completion and better output now.

Providers of instructional AI should be running randomised controlled trials and publishing them where practitioners can find and compare them. Stanford’s SCALE Initiative maintains an AI for education research repository for exactly this. A claim about learning gains that rests on engagement metrics and completion rates is not a claim about learning.

Assessment is already moving without waiting for us. In May, Princeton’s faculty voted to proctor all in-person examinations from July, ending a practice that had stood since 1893. The proposal named AI and small personal devices directly, noting they have made misconduct much harder for other students to observe. Institutions are rebuilding the places where they can still see unassisted performance. Products that cannot show anything about unassisted performance will have less and less to say to them.

The problem is not technological

The authors of the China study end somewhere uncomfortable for builders. Well designed tutoring tools already exist, in China and in many other countries, often at nearly zero cost. Most students do not use them. They prefer general purpose tools that make it easy to hand over the homework. A separate two-year school experiment found that students given access to an AI tutor configured to coach rather than answer simply did not engage with it.

Better models will not fix that. The essential question in AI adoption in education is organisational and institutional rather than technological. It is about what schools count, what teachers can see, what students are rewarded for, and what the work is actually for.

For those of us building, the immediate move is narrower and entirely within reach. Stop shipping the homework score as a proxy for learning. Start collecting something that still means what we think it means.

Sources

Strömberg, D., Lei, V., and Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. CEPR Discussion Paper 21577. https://cepr.org/publications/dp21577

Same paper on SSRN, DOI 10.2139/ssrn.6868618. https://ssrn.com/abstract=6868618

Strömberg, D., Lei, V., and Wu, Y. (2026). The generative AI learning penalty in secondary school. VoxEU, 4 September 2026. https://cepr.org/voxeu/columns/generative-ai-learning-penalty-secondary-school

OECD (2026). PISA 2025 Results (Volume I): Future-Ready Students. https://www.oecd.org/en/publications/pisa-2025-results-volume-i_73451bc5-en.html

Bastani, H. et al. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. PNAS 122(26).

Kestin, G. et al. (2025). AI tutoring outperforms in-class active learning. Scientific Reports 15: 17458.

Oreopoulos, P. and Low, N. (2026). One click away: AI tutoring with Khanmigo in a two-year school experiment. NBER Working Paper 35620.

The Daily Princetonian (2026). Princeton faculty mandate proctoring for in-person exams, upending 133 years of precedent. 11 May 2026. https://www.dailyprincetonian.com/article/2026/05/princeton-news-adpol-proctoring-in-person-examinations-passed-faculty-133-years-precedent

Stanford SCALE Initiative AI for education research repository. https://scale.stanford.edu/ai/repository/all?study_design%5B55%5D=55

📬 Want more insights like this?

Subscribe to Grounding EdTech and get weekly insights on AI, EdTech, and instructional design — plus free access to our Instructional Design for Educators course.

No spam. Unsubscribe anytime.

Leave a Reply

Your email address will not be published. Required fields are marked *