Photo: Unsplash
Standardized Tests' Last Stand
The College Board’s revenue in fiscal year 2023 was $1.94 billion. The SAT is the largest revenue generator. The system that feeds into it, test prep courses, tutoring, Khan Academy’s SAT prep partnership, score-sending services, represents several billion more. Standardized testing is not just a pedagogical instrument; it is an industry, with constituencies, a political infrastructure, and institutional survival incentives that operate independently of its educational merit.
This matters for understanding what’s happening as AI destabilizes the assessment ecosystem. The standardized testing industry is not primarily asking “how should we assess students in a world with AI?” It’s asking “how do we maintain relevance and revenue while AI changes what we can credibly measure?” These are different questions with different answers.
The SAT has survived two previous technology threats, and both survivals are instructive. The calculator threat in the 1970s-1990s was handled by eventually allowing calculators on the math section while redesigning the hardest problems to require conceptual understanding that calculators don’t supply. The problem reformulation was real. The calculator-era SAT math questions emphasize reasoning about functions, identifying relevant information, and interpreting results, rather than pure computation. The test adapted.
The internet threat in the early 2000s was different in kind. The concern was that test content would leak online, allowing students who found leaked questions to outperform students who hadn’t. The College Board’s response was to move toward adaptive digital testing, question banks large enough that any particular student’s test contained items unlikely to have been leaked, and eventually the digital SAT launched in 2023, which uses a computerized adaptive format that selects questions in real time based on student performance.
Both adaptations worked, in the sense that the SAT remained a credible-enough signal to maintain market position. Neither produced a fundamentally better assessment. Both were defensive moves to preserve the test’s validity against specific threats. The underlying construct being measured, a blend of verbal reasoning and mathematical ability, measurable in 2-3 hours under standardized conditions, remained essentially unchanged.
AI presents a different structural challenge. The previous threats were about test security (content leaking) and test accessibility (calculators giving an edge). AI threatens the construct validity of timed, written assessments in a more fundamental way. The SAT Reading and Writing section measures whether students can analyze text, identify the central idea, evaluate argument structure, and make inferences. These are legitimate cognitive skills. They also happen to be skills that GPT-4-level AI can demonstrate at high levels under test conditions.
The College Board has responded to this by implementing its own AI detection on written components, building test conditions that restrict device access, and moving the essay component to an optional add-on that is taken under controlled conditions. These are reasonable defensive measures. They don’t address the underlying question: if the skills the SAT measures are skills that AI can demonstrate, what does a high SAT score certify about a student in a world where AI is universally available?
The answer the College Board would give is: it certifies that the student has those skills independent of AI assistance. This is valuable, because there are contexts where students need to reason without AI access. A college seminar that requires real-time intellectual engagement, a professional environment where decisions must be made quickly without a device: in these contexts, having the underlying skills matters.
This argument is defensible. But its strength depends on how often those contexts arise for the students who are being assessed, and the frequency is declining.
The testing companies are not wrong to defend the human, unassisted baseline. The question is whether the SAT is the right instrument for measuring it, and whether the 3-hour standardized test format is still the most efficient and fair way to identify students who have the underlying capabilities colleges care about.
Colleges have been asking this question for a while. The COVID-19 pandemic forced large-scale test-optional admissions starting in 2020, and the test-optional experiment ran for 3-4 years at hundreds of institutions. The results are contested — some studies found that test scores remained predictive of college GPA even after controlling for other factors, while others found that other application components were equally or more predictive. The Massachusetts Institute of Technology reinstated mandatory SAT requirements in 2023, citing internal data showing that test scores added predictive value that other application components didn’t capture. Yale and Dartmouth followed. Harvard has maintained test-optional but studies its own predictive data carefully.
The evidence is that standardized test scores continue to predict college academic performance, even in an AI era. This is not surprising. The cognitive skills they measure are real, and real skills predict real performance. The more interesting question is whether those skills could be measured more efficiently with different instruments, particularly instruments that are harder for AI to game.
The most interesting development in standardized testing as of 2025-2026 is not AI detection. It’s performance-based assessment at scale. Several universities have experimented with portfolio-based admissions in which students submit evidence of their work process, not just final products. Some have implemented virtual interviews that probe reasoning in real time. A few graduate programs have moved to “AI-transparent” admissions in which applicants are told they may use AI assistance during application and are assessed on the quality of their judgment about when and how to use it.
These approaches are more expensive to evaluate than a scan-scored standardized test. They’re also harder to fake. A student who submits a research paper showing thoughtful iteration and genuine engagement is demonstrating something different from a student who submits a polished AI-generated product, and a skilled admissions reader can usually tell the difference, particularly when combined with other signals.
The standardized testing industry’s most important adaptation will be the one that’s hardest. Redesigning what they test, not just how they prevent cheating. A test designed for 2030 would measure the skills that remain valuable when AI is available as a tool: the ability to evaluate the quality of an AI output, to identify when AI reasoning has gone wrong, to make judgment calls about ambiguous situations that don’t have single correct answers, to integrate information from multiple conflicting sources, to make defensible decisions under uncertainty.
These are harder to test in a standardized format. They require more context, more judgment in grading, more design investment. They’re less amenable to the machine-scored formats that make large-scale testing economically viable. But they’re the skills that actually matter in the world the students will graduate into.
The College Board’s 2024 annual report mentioned AI 23 times. None of the mentions described a plan to fundamentally redesign the construct being measured. They described plans to maintain security, add AI literacy components, and expand the digital format. This is the industry responding to AI as a threat to manage rather than as a forcing function to rethink the purpose of the assessment.
Standardized tests may well survive as a useful signal in a form recognizable to current users. But the industry’s long-term relevance depends on whether it can make the harder argument, that what it measures is genuinely important in a world with AI, clearly enough and redesign its products to match that argument. So far, the argument is being made in retrospect rather than prospectively, and the product is being defended rather than reconceived.
The ACT’s trajectory is instructive. The ACT lost market share to the SAT starting in the mid-2010s, partly because College Board invested aggressively in partnerships, the Khan Academy free SAT prep being the most prominent, that made the SAT more accessible and better-known among students who had previously defaulted to the ACT. By 2023, SAT participation was three times ACT participation nationally, reversing decades of rough parity. The ACT responded by launching a “superscoring” option (using the best section scores across multiple sittings), a shorter version of the test, and online administration options. These were competitive responses, not pedagogical ones.
The market dynamics of standardized testing will continue to push toward feature competition — more flexible administration, lower cost, more institutional acceptance — rather than toward the harder work of redesigning what gets measured. New entrants like Duolingo’s English Test, which established itself as a cheaper and more accessible alternative to the TOEFL for international student English proficiency assessment, show that the incumbent testing giants can be disrupted from below. But Duolingo’s test is still measuring the same construct as the TOEFL. It just measures it more efficiently and cheaply.
The test that would genuinely win in an AI-disrupted higher education environment would measure the skills colleges and employers actually want: the ability to identify the right question to ask, to synthesize incomplete information under time pressure, to explain complex ideas to different audiences, to change your mind when evidence changes. These are difficult to standardize, expensive to grade, and don’t fit into the 3-hour bubble-sheet format. They also can’t easily be gamed by AI — a student who needs to generate, evaluate, and defend a novel solution to a genuinely ambiguous problem is demonstrating something that requires genuine cognitive engagement, which is exactly what AI can’t be handed to fake for them.
Whether the testing industry builds this test depends on whether the colleges and employers who are the customers of testing data are willing to demand it. So far, the demand has been weak. The status quo is comfortable for colleges that have built admissions processes around existing scores. Disrupting the measurement instrument means disrupting the admissions process, which means disrupting the implicit social contract about what college selectivity is for. That’s a much bigger fight than the testing companies want to pick.
The trajectory is one of slow erosion rather than sudden collapse. Standardized tests will not disappear — they’re too deeply embedded in college admissions infrastructure, financial aid allocation, and scholarship programs for any rapid divestment. But their role will gradually narrow. Highly selective institutions will use them as one signal among several in increasingly holistic processes. Employer use will continue declining as alternative assessment options proliferate. The fraction of admissions decisions where test scores are the primary differentiator will shrink.
What survives of standardized testing will probably be more specific: not broad aptitude tests but assessments of specific domain knowledge for students entering competitive programs where baseline knowledge genuinely matters: medical school prerequisites, law school, graduate engineering. These applications have always been the strongest use case for standardized assessments, because the relevant baseline competencies are real and consequential. The broad college admissions use case, measuring “general readiness,” is where the credibility has been most eroded, and where the AI disruption will bite hardest.
The industry that navigates this well will be the one that contracts intelligently rather than fighting every contraction. That requires a clarity about where standardized testing genuinely adds value that the current leadership of the College Board and ACT hasn’t yet demonstrated.
One email a month: the upcoming live event + free recording access for subscribers. No spam, unsubscribe anytime.



