A student uses AI to finish homework faster and receives a better grade.
Then the AI is removed and the student performs worse.
That sounds like an argument for banning AI from education. It is not. It is an argument for being much more precise about what we think an assessment measures.
In June 2026, David Strömberg, Victor Lei and Yanhui Wu published a CEPR discussion paper based on 30 months of data from 26,811 Chinese students. After students adopted generative AI, homework scores rose by 18% and completion time fell by 30%. Within six months, monthly closed-book exam scores fell by 20%. The losses were concentrated among the roughly 80% of AI users whose unusually fast, high-scoring homework suggested that they were outsourcing the work.

Chart published by The Economist, based on Strömberg, Lei and Wu's CEPR discussion paper.
This is serious evidence, but it is not the final word. The CEPR paper is a discussion paper, not a peer-reviewed randomized trial. Its difference-in-differences design can identify a strong pattern in a large real-world dataset, but it cannot prove that every student, subject or education system will respond in the same way.
A separate experiment makes the mechanism harder to dismiss. In a peer-reviewed randomized controlled trial published in PNAS, nearly 1,000 Turkish high-school students used one of two GPT-4 systems while practising mathematics. A standard ChatGPT-like interface improved practice performance by 48%, yet those students scored 17% worse than the control group when the tool was removed. A safeguarded tutor that gave hints rather than answers improved practice performance by 127% and largely eliminated the later penalty.
The important variable was not simply whether AI was present.
It was what the AI allowed the student to avoid doing.
The exam did not fail
It is tempting to look at these results and say that closed-book exams are relics of an information-scarce age. In most professional settings, people can search, calculate, collaborate and use software. Increasingly, they can also use AI. An examination hall that removes every tool does not reproduce the conditions in which graduates will work.
But in these studies, the unaided exam performed an essential function: it exposed the difference between producing an answer and developing the competence to produce, inspect and defend one.
If the exam had disappeared, the homework grades would have told a reassuring but false story. Students were faster. Their answers looked better. The output metric improved while the underlying capability deteriorated.
So the correct conclusion is not that exams are obsolete.
It is that a one-mode assessment system is obsolete.
Traditional exams usually measure only what a student can do alone, from memory, under time pressure. Unsupervised assignments increasingly measure what a student and an undisclosed collection of tools can produce together. Neither result is sufficient on its own.
Education now needs to measure three things separately:
- What can the student do without assistance?
- What can the student accomplish with AI?
- Can the student move responsibly between those two modes?
Mode one: independent competence
Some knowledge must be available without a prompt, search result or generated explanation.
You cannot evaluate an AI answer in a domain you do not understand at all. You cannot notice a broken assumption, an impossible unit, a fabricated source or an elegant argument built on a false premise if you have no internal model against which to test it.
Independent assessment should therefore remain—but become narrower and more purposeful. Short supervised checks can test foundational knowledge, mental models, reasoning steps and transfer to unfamiliar problems. The goal is not to reward the largest store of memorized facts. It is to establish the minimum competence required to think, learn and verify without the tool.
The distinction matters. Memorization is not the same as mastery, but mastery is impossible when every relevant concept must first be retrieved from somewhere else.
Mode two: tool-assisted judgment
The other half of modern competence appears only when the tools are available.
A student working with AI should be assessed on whether they can:
- frame a useful problem rather than merely request an answer;
- choose when AI is appropriate and when it is not;
- inspect claims against primary sources and other evidence;
- recognize uncertainty, bias and missing context;
- compare alternatives instead of accepting the first plausible output;
- improve a weak response through iteration;
- disclose significant AI contributions;
- and remain accountable for the final work.
These are not shortcuts around learning. They are increasingly part of the work that learning is meant to prepare people to do.
The OECD's 2025 review of AI adoption in education points in this direction. Its planned PISA 2029 Media and Artificial Intelligence Literacy assessment uses simulated digital and AI tools for tasks such as checking the accuracy of social-media claims and evaluating privacy policies. That is a different kind of test: the tool is part of the environment, and judgment is the capability being measured.
Mode three: the bridge
The most revealing assessment may be the transition between independent and assisted work.
Ask a student to explain why they accepted one AI-generated claim and rejected another. Change a constraint and see whether they can adapt the solution. Remove the tool for one critical step. Present an error and ask them to diagnose it. Require an oral defence of the final artifact and the process that produced it.
This bridge makes intellectual ownership visible.
An AI-generated project can look excellent even when the student understands very little. A short conversation about the assumptions, trade-offs and discarded alternatives is much harder to outsource. The OECD highlights unscripted oral assessment as both resistant to inappropriate AI assistance and useful for testing critical reasoning and adaptive response. UNESCO's assessment work similarly argues for evaluating process, dialogue, defence and authentic problems rather than relying only on final artifacts.
The bridge also changes the incentive. If students know they may be asked to defend any part of their work, using AI to bypass understanding becomes a risky strategy. Using it to explore, challenge and improve understanding becomes valuable.
What a redesigned course could look like
There is no universal percentage split. Mathematics, history, medicine, design and computer science do not need identical assessment systems. But a course could combine:
- Brief closed-tool checks for foundational knowledge and unaided reasoning
- AI-permitted projects based on authentic, open-ended problems
- Process evidence such as drafts, source trails, decisions and significant prompts
- Oral or live defences that test ownership and adaptive reasoning
- Reflection after tool removal to reveal what transferred from assisted practice into durable competence
The Australian Tertiary Education Quality and Standards Agency's assessment-reform guidance makes a similar systemic point: institutions need program-level assessment that preserves trustworthy evidence of learning while preparing students to participate ethically in a society where AI is normal.
This design is more demanding than either banning AI or allowing it everywhere. It requires teachers to specify what is being assessed, provide equitable access to permitted tools, protect student data, define disclosure rules, and design tasks whose value survives model improvement. Oral defences and process reviews also require time.
But the administrative convenience of the old exam should not be confused with educational validity.
Stop grading the artifact as if it were the person
Generative AI has made a long-standing measurement problem impossible to ignore.
We often grade the artifact—a homework solution, essay, report or piece of code—and treat it as evidence of the person. That inference was never perfect. Tutors, parents, peers, templates and internet search already blurred the boundary. Generative AI widened the gap dramatically.
The answer is not to build a better detector for the artifact. It is to collect better evidence about the learner.
Can the student reason without assistance?
Can they produce better work with assistance?
Can they explain where the tool helped, where it failed and why the final decision remains theirs?
Those are different capabilities. A serious assessment system should make all three visible.
The old examination hall should lose its monopoly, not disappear. It remains useful precisely because independent competence still matters. What no longer makes sense is pretending that unaided performance is the only legitimate form of intelligence—or that polished AI-assisted output proves learning.
The future belongs to people who can think without AI and think better with it.
Our exams should be able to tell the difference.
Sources
- Strömberg, Lei and Wu's CEPR discussion paper provides the 30-month panel evidence on homework productivity and later closed-book performance among 26,811 Chinese students.
- Bastani et al.'s PNAS randomized controlled trial compares an unrestricted GPT-4 interface, a safeguarded tutor and a no-AI control in high-school mathematics.
- The OECD report on AI adoption in education discusses oral assessment, curriculum reform, tool-assisted AI literacy and the planned PISA 2029 assessment.
- TEQSA's assessment-reform guidance sets out program-level directions for trustworthy learning evidence and ethical participation with AI.
- UNESCO's IdeasLAB essay “What's worth measuring? The future of assessment in the AI age” develops process-, dialogue- and authentic-problem-based alternatives to artifact-only assessment.
- The Economist's August 18 chart and analysis prompted this essay and visualizes the CEPR findings.
