Photo: Unsplash
What We Talk About When We Talk About Critical Thinking
Ask ten school administrators what skill becomes most important in an AI-saturated world, and nine of them will say “critical thinking.” Press any of them to define critical thinking precisely enough to build a curriculum around it, and you will get a range of answers that are, on close examination, different things. Evaluating sources, logical reasoning, creative problem-solving, metacognition, systemic thinking, questioning assumptions. These are real skills. They are not the same skill.
The failure to be precise about what critical thinking means is not an academic problem. It’s a practical disaster. You cannot teach “critical thinking” as an undifferentiated concept. You can teach students to construct a logical argument and identify the premises. You can teach them to evaluate the credibility of a source based on its methodology and track record. You can teach them to notice when an analogy breaks down. You can teach them to identify unstated assumptions in a position they’re being asked to accept. All of these are better than teaching none of them. None of them are the same thing, and treating them as a single skill has produced decades of curricula that gesture toward critical thinking while teaching specific skills inconsistently and assessing them almost not at all.
AI makes this crisis acute because AI can produce text that has the surface features of critical thinking, organized structure, logical connectives, qualifications and counterarguments, while containing very little of the substance. An AI-generated essay on whether social media is harmful will identify multiple perspectives, acknowledge counterarguments, and reach a measured conclusion. It will also be entirely derivative, containing no original analysis, drawing on no genuine engagement with specific evidence, and requiring from the student zero cognitive effort beyond framing the prompt. When students submit this and pass, schools have revealed that they were measuring the surface features, not the substance.
The surface features were always a noisy proxy for the substance. We tested whether a student could write a correctly structured argument because correctly structuring an argument is evidence that you had to organize your thinking. AI has eliminated the inferential step. Now the structure is available without the thinking, and we have to measure the thinking directly.
This should not be hard, in principle. You can ask a student to defend their argument orally. You can ask them to explain why they chose one piece of evidence over another. You can ask them to identify the strongest objection to their position and explain why it doesn’t defeat their argument. You can give them a new piece of evidence and ask whether it changes their conclusion. All of these require genuine engagement with the ideas in a way that’s very difficult to outsource to AI in real time.
The problem is that these assessments are expensive. An oral examination of a student’s reasoning takes 15-20 minutes of a teacher’s time, one-on-one, and requires the teacher to be able to engage substantively with whatever the student argues. In a class of 30 students, that’s 7-10 hours of individual oral assessment per assignment. This is approximately what Oxbridge tutorial systems have done for centuries, and the educational outcomes are consistently excellent. It requires a ratio of about 1 tutor to 6-8 students and is economically feasible only at institutions with substantial resources and centuries of tradition.
The solution AI actually offers here is not that it replaces the teacher in conducting critical thinking assessments. It’s that it can help design them more cheaply. An AI can generate targeted follow-up questions calibrated to a student’s actual argument. It can run a Socratic dialogue that pushes back specifically on the claims the student has made. It can flag when a student’s response is evasive, when the reasoning contains a gap, when the counterargument presented is a strawman. This doesn’t fully replace the human examiner, but it reduces the human time required and can scale oral-style assessment to more students than purely human evaluation can.
There’s a deeper problem with the critical thinking agenda that the AI-in-education conversation hasn’t confronted: critical thinking is domain-specific to a much greater extent than the rhetoric suggests.
This is a finding from cognitive science that has been settled for decades and that educational policy has largely failed to absorb. The general critical thinking skills that transfer across domains are much narrower than the skills that are specifically useful within a domain. A person who can construct a rigorous argument about constitutional law cannot necessarily construct a rigorous argument about statistical methodology. A scientist trained to evaluate evidence rigorously in their field applies that rigor systematically within their field and imposes it inconsistently outside it.
The implication for curriculum design is that you develop critical thinking by teaching students to think critically within specific domains, teaching them the standards of evidence and argument in history, in science, in mathematics, in literature, not by trying to teach “critical thinking” as a subject that somehow generalizes. The subjects matter. The discipline-specific epistemologies matter. You can’t route around them with a generic thinking skills class.
AI creates pressure toward exactly this error. When “AI handles the content” and “students develop critical thinking,” the implicit model is that critical thinking is separable from domain knowledge. It isn’t. A student who knows nothing about climate science cannot critically evaluate a climate claim, because they can’t distinguish a sophisticated-sounding claim that violates physical principles from a sophisticated-sounding claim that doesn’t. Background knowledge is the prerequisite to critical evaluation, not an alternative to it.
The specific skill that AI genuinely makes more important, that is currently underdeveloped in most curricula and that deserves the “critical thinking” label more than most things assigned that label, is epistemic calibration. Knowing what you know, knowing what you don’t know, and knowing the difference between a well-supported belief and a plausible-sounding one.
AI models are systematically overconfident. They produce plausible text with high fluency regardless of whether the underlying claims are well-supported. A student who cannot tell the difference between a confident, well-reasoned argument and a confident, fluent non-argument is dangerously exposed to manipulation by AI-generated content. This isn’t a new skill — being suspicious of confident claims has always been a useful disposition — but AI makes it more consequential and more difficult to exercise, because the surface quality of AI-generated text is much higher than the surface quality of low-quality human writing. Bad human arguments usually read poorly. Bad AI arguments can read brilliantly.
Teaching students to attend to the quality of evidence and reasoning rather than the quality of prose, to ask “what is this argument’s evidence?” rather than “is this well-written?”, is specific, teachable, and genuinely more important in an AI-saturated world than it was before. It also requires that students have enough background knowledge in the relevant domain to evaluate the evidence, which brings us back to why domain knowledge remains important.
The version of critical thinking education that AI makes obsolete is the kind built around producing polished analytical prose — the five-paragraph essay, the thesis-driven argument paper, the structured analytical report. These forms were useful partly as epistemic scaffolding, forcing students to structure their thinking, and partly as output assessments, measuring whether students could produce organized, evidence-supported arguments. AI can now produce these forms easily, which makes them poor assessments. But the epistemic scaffolding function was the valuable one, and it can be preserved in forms that AI can’t substitute: live dialogue, real-time reasoning, collaborative argument under pressure.
The schools that figure this out will redesign their approach to critical thinking from “produce an argument” to “defend and develop an argument in real time.” That shift requires different assessment methods, different classroom structures, and a higher baseline of teacher expertise. It’s harder and more expensive than the current model. It’s also what education has always been at its best: the experienced reasoner challenging the student’s thinking in real time, not accepting the well-formed surface but probing the understanding underneath.
There’s a specific skill that deserves its own mention because it’s been almost entirely absent from educational curricula and is now urgently needed: probabilistic reasoning about confidence. Students, and most adults, have almost no training in the difference between “I believe this confidently” and “the evidence supports this strongly.” These are related but distinct claims. A person can be genuinely confident about a false belief. Evidence can strongly support a conclusion that is, in retrospect, wrong. The failure to maintain this distinction is responsible for an enormous amount of bad decision-making, from individuals all the way to national governments.
AI makes this worse in a specific way. LLMs are trained to produce fluent, confident-sounding text. The model’s “confidence” in the technical sense, the probability distribution over next tokens, doesn’t track epistemic justification. The AI sounds equally confident when it’s drawing on well-established facts and when it’s confabulating. The student who hasn’t been trained to ask “what’s the evidence for this, and how strong is it?” is poorly equipped to use AI safely, because they have no mechanism for distinguishing the reliable outputs from the unreliable ones beyond surface plausibility.
Teaching probabilistic reasoning about confidence is not glamorous. It doesn’t produce the kind of one-line curricular objective that appears in state standards. It requires practice across many domains and contexts, with feedback from people who can model appropriate calibration. It’s also more important than almost anything else listed under “critical thinking” in most school curricula. An 18-year-old who knows how to ask “how confident should I be in this, and why?” is better equipped to be a citizen and a professional than one who can structure a five-paragraph argument but accepts AI-generated conclusions uncritically.
The AI-era curriculum that actually prepares students for the world they’re entering looks less like the current model than most schools are willing to acknowledge. It involves more live dialogue and less essay production. More domain-specific epistemology and less generic “critical thinking skills.” More genuine intellectual difficulty and less frictionless answer retrieval. More explicit attention to what it feels like to not know something and to be working toward knowing it. The technology exists to support this kind of education at scale. What remains is the institutional will to build it.
One email a month: the upcoming live event + free recording access for subscribers. No spam, unsubscribe anytime.

