Photo: Unsplash
Why Schools Are Getting AI Wrong
In January 2023, New York City public schools blocked ChatGPT on school networks. Within about fifteen months Turnitin was running AI detection across tens of millions of submitted papers, and most large universities had rewritten their integrity policies to name large language models specifically.
All of it was the wrong move. Not tactically wrong. Wrong at the level of what problem anyone thought they were solving.
When a technology changes what a student can produce, the useful response is to rethink what you’re measuring and why you’re measuring it. Banning the thing, or trying to detect it, delays the reckoning without preventing it. Every school that spent 2023 and 2024 refining its detection policy was fighting the last war. The kids who learned to use these tools well during those two years now have an advantage over the kids who were successfully policed into not using them.
Everyone reaches for calculators as the parallel, and it is the right parallel, just not for the reason people cite. The usual version goes: we let calculators in, math education survived, students still learned math, so this will be fine too. Trust the process.
That reading is far too comfortable. Calculators did not arrive smoothly. The National Council of Teachers of Mathematics recommended their use in 1974. Districts resisted through the late seventies. The SAT didn’t permit them until 1994, twenty years after the recommendation, and the ACT waited until 1996. What happened in the gap is that a generation of teachers kept drilling arithmetic that was already becoming economically worthless, while nobody built the curriculum that would have used the calculator to reach the things that are genuinely hard without one. Graphs. Conjectures. Ugly real-world data sets with missing rows.
By the time the tests caught up, that whole layer of mathematics had barely been touched, because the old curriculum left no room for it.
The schools that handled calculators well did not “allow” them. They rewrote what the subject was for. Instead of checking whether a fifteen-year-old could multiply three-digit numbers by hand, they checked whether she could build a model, reason about a distribution, notice when an answer was absurd. The technology shift was the occasion for getting honest about the point of the class.
Most schools have not started that conversation about AI. They are still at the policy-writing stage, and the policies don’t hold. Turnitin’s accuracy claims have been contested from the start: in 2023 a Stanford group ran seven detectors over TOEFL essays by non-native English speakers, and more than half came back flagged as machine-written, while essays by American eighth-graders passed cleanly. That isn’t a rounding error. That’s one specific group of students getting accused for writing in a second language. Meanwhile the evasion tricks circulate on Reddit and TikTok faster than the vendors ship updates.
This arms race ends the way spam filtering did. The detector defines the frontier and the generator walks around it.
Treating any of this as a winnable enforcement problem is self-deception.
What makes the situation hard is that the case for enforcement isn’t stupid. It rests on something real: if a student submits a generated essay, what did they learn? Probably less than if they’d written it. Retrieval practice, the business of actively reconstructing what you know instead of re-reading it, produces better retention than passive exposure, and that result has held up for decades. Writing an essay is hard partly because it forces you to organize what you know, find the holes, and build an argument. Hand that to a model and most of the friction goes away.
The concern is legitimate. The conclusion drawn from it is not.
Put two students side by side. The first runs a well-designed AI tutoring session: she gets asked questions, discovers what she can’t explain, receives an explanation aimed at her specific gap, and defends an answer under something like Socratic pressure. Cognitive load high. The second pastes a prompt, copies the output, submits it. Cognitive load roughly zero.
Same underlying technology. A policy that treats those two identically, on an allowed-or-not axis, is making a category error. The line that matters isn’t AI-assisted versus not. It’s between work that requires thinking and work that can be handed off without any. So the curriculum question was never “how do we detect this.” It’s “how do we set tasks that can’t be finished without the student actually thinking.”
That question is much harder than writing a better honor code, which is roughly why nobody has answered it.
There’s a class dimension to this that stays almost entirely out of the official conversation. These tools are not evenly distributed. A student with reliable internet, a laptop that works, and a parent paying for the good tier has help that a student without those things doesn’t. When schools in wealthy districts quietly tolerate AI use while under-resourced districts try to enforce a ban, the students with the least help at home are the ones getting disciplined for finding some.
I don’t have a clean number for the size of that gap, and I distrust most of the ones I’ve seen quoted, including the ones that support my argument. What I do trust is the mechanism, because it’s the same one that showed up with home internet in the 2000s and with private tutoring before that: an advantage that starts as access ends up encoded as merit.
Underneath all of it, schools don’t agree on what they’re trying to accomplish, and AI has dragged that disagreement into the open.
If the point of an essay assignment is to check whether a student can produce grammatical prose organized into paragraphs, then yes, the assessment is broken, because that skill no longer separates anyone from anyone. If the point is to find out whether a student can think carefully about something, build an original argument, and hold it under interrogation, then the essay was always a proxy and a noisy one. AI didn’t break the goal. It broke a measuring instrument that was already bent.
The conversation schools need to have out loud: which skills are we developing, and which of those still matter when everyone has a model in their pocket? Writing a clean five-paragraph essay isn’t on the list. Evaluating an argument, spotting the gap in it, defending a claim while someone pushes back, revising a model when the data won’t cooperate: those are. And the awkward part is that AI, used properly, is a better instrument for building those skills than the pre-AI curriculum ever was. You can have a model argue against every claim you make, in real time, pitched at your level, instead of waiting two weeks for a paper to come back with a B+ and no comments.
The schools that work this out first will look back at the detection era the way we look back at the districts that banned calculators in the eighties.
One more thing worth naming: the detection industry’s own epistemics. Turnitin, GPTZero, Copyleaks and the rest have a financial interest in the problem staying unsolved. Every flagged paper validates the product. Every policy update that requires a detector is a renewal. The whole category rests on a definition of the problem, AI use is cheating, that happens to be the definition that sells licenses. When a detection vendor speaks at an education conference, they are not a neutral party advising schools on their interests.
A school that makes detection central to its integrity policy has handed policy design to a supplier. That’s not new in education. Textbook publishers have shaped curriculum for decades and testing companies have shaped what gets taught since the seventies. But this case is sharper, because the vendors are selling a fix for a contested problem, using accuracy claims that are disputed, in a way that bends institutional incentives against the redesign that would actually help.
And the institutional gravity runs deeper than procurement. A university that rebuilds its assessments to be AI-proof no longer needs detection software. It also no longer needs most of the apparatus that grew up around Turnitin: the integrity officers, the honor councils, the adjudication procedures. Those are jobs. Redesigning assessment threatens them. Buying software doesn’t.
So what would a defensible policy look like? Roughly this. Say clearly where AI help is expected, in research, drafting, project work, and where it isn’t, in class assessments, oral exams, live demonstrations of skill. Put money into assessment that requires engagement: vivas, portfolios with visible revision history, group projects with individual accountability. Teach the tools as a subject, including where they fail. And drop the fiction that an honor code can outrun a technology.
None of that is radical. It’s what good assessment design looked like before any of this, adapted to the fact that the machines on the other side of the desk changed. The difficulty is organizational, not intellectual. Schools aren’t getting this wrong because they don’t understand it. They’re getting it wrong because the parents want education that looks familiar, the administrators want a policy that can be enforced, and a vendor is standing in the hallway with a quote.
One email a month: the upcoming live event + free recording access for subscribers. No spam, unsubscribe anytime.





