The democratization narrative around medical AI has always had a certain utopian logic. In a world where cardiologists are concentrated in wealthy urban centers and rural communities struggle to attract any specialists at all, an AI that can read an ECG with cardiologist-level accuracy and run on commodity hardware represents something genuinely valuable: the beginning of expert-level diagnosis available anywhere. The same argument applies to dermatology AI (a skilled dermoscopy AI in a primary care office), radiology AI (a chest X-ray reader that flags critical findings in a hospital without a 24-hour radiologist), and pathology AI (slide grading that doesn’t require subspecialist review).

The argument is logically sound. The deployment reality is considerably grimmer.

Where AI Diagnostics Are Actually Deployed

A 2024 study by the Brookings Institution analyzed the deployment geography of FDA-cleared AI diagnostic tools across U.S. health systems and found a pattern that should surprise nobody who has watched healthcare technology adoption: the most AI-equipped hospitals are concentrated in urban areas, affiliated with academic medical centers, and serving higher-income patient populations. Rural hospitals, community health centers serving predominantly Medicaid populations, and federally qualified health centers — the institutions closest to the patients who lack access to specialists — have the lowest rates of AI diagnostic deployment.

This is not primarily a technology access story. The tools themselves are not restricted to wealthy institutions. It is an economics and infrastructure story. AI diagnostic tools require integration with electronic health record systems, imaging infrastructure, IT support capacity, and procurement budgets that community hospitals often lack. The contractual and implementation complexity of deploying even a single AI diagnostic tool is substantial, and many smaller institutions lack the administrative capacity to manage it. The result is that AI follows existing health system resources rather than filling health system gaps.

The Training Data Problem Revisited

Even where AI tools are deployed in underserved settings, the performance disparities documented in several published studies are significant. The underlying issue is training data composition. Most AI diagnostic models were trained predominantly on data from academic medical centers with specific demographic compositions. When those models are deployed to patient populations that differ demographically from the training set, performance degrades.

The most extensively documented case is dermatology AI. Several widely used AI skin lesion classifiers show materially lower sensitivity for melanoma in patients with darker skin tones. A 2023 review in JAMA Dermatology found that of 33 clinical studies of AI dermatology tools, only 12 reported model performance by Fitzpatrick skin type, and among those that did, 8 showed a meaningful performance gap. The gap is partially a consequence of training data: most skin lesion datasets are composed predominantly of images from patients with lighter skin, which is a direct reflection of the populations historically included in dermatology research.

This is a concrete harm in a concrete clinical setting. A patient with darker skin, in an underserved community without easy access to a dermatologist, using an AI screening tool, may receive a false reassurance about a lesion that the tool was never reliably calibrated to evaluate. The tool was supposed to reduce the access gap. Instead it introduced a quality gap that layered on top of the access gap.

Dermatology is not the only affected specialty. Pulse oximetry AI and algorithms have known performance gaps in patients with darker skin. Cardiovascular risk prediction algorithms trained on datasets that underrepresent certain ethnic groups produce systematically miscalibrated predictions when applied to those populations. The depth and breadth of training data bias in medical AI is extensive, and much of it is only partially characterized even in the published literature.

The Feedback Loop Problem

There’s a structural reason why this problem is hard to fix and why well-intentioned efforts to address it often fall short. Improving model performance on underrepresented populations requires more labeled data from those populations. Getting labeled data requires the cooperation of health systems serving those populations. Health systems serving underrepresented populations tend to have less research infrastructure, fewer data science resources, and less incentive to participate in data-sharing agreements that primarily benefit commercial AI vendors.

The AI vendor profits most from data collected from high-volume, well-resourced institutions that process many patients efficiently. The ethical imperative to include underrepresented populations in training data is not, structurally, well-aligned with the commercial incentive to collect the most data most efficiently. NIH, AHRQ, and several philanthropic organizations have attempted to fund programs that close this gap, with partial success. The scale of the effort required exceeds what research funding alone can provide.

FDA guidance issued in 2024 required AI medical device submissions to include subgroup performance analyses and describe the demographic composition of validation datasets. This is a meaningful step. It doesn’t retroactively fix the models already deployed, and it doesn’t address the gap between reporting performance disparities (required) and actually achieving equity (not yet required).

The Algorithmic Triage Problem

Beyond diagnostic AI, a related and underappreciated problem is algorithmic triage and resource allocation in healthcare systems. Hospital sepsis prediction algorithms, ICU deterioration models, and emergency department severity scores are now routinely deployed in large health systems to help allocate nursing attention, predict who needs intensive monitoring, and flag patients likely to deteriorate.

A 2023 study in Nature Medicine evaluated a widely used commercial sepsis prediction algorithm (Sepsis-3 compliant, deployed in multiple major health systems) and found that Black patients were significantly less likely to be flagged as high-risk at the same clinical condition severity as white patients. The algorithm, trained on historical outcomes data, was learning patterns that reflected not just disease severity but historical differences in how care was allocated to different patient populations — differences that may have included systematic undertreatment of Black patients in the training data’s historical period. The algorithm was encoding historical inequity as if it were biological signal.

This is the mechanism by which AI can actively amplify inequity rather than merely failing to reduce it: if training data reflects historical practices that were themselves inequitable, and the model learns to predict outcomes in that training data, it will perpetuate those outcomes in its predictions. This is not a fringe theoretical concern — it is a documented phenomenon in deployed systems.

The Path That Actually Exists

None of this means AI diagnostics are net harmful or should not be deployed. The evidence that these tools improve care in the settings where they work well is real. The case for deploying chest X-ray triage AI in an emergency department that doesn’t have 24-hour radiologist coverage is compelling. The case for using diabetic retinopathy screening AI in primary care practices in underserved communities is strong.

The honest framing is that AI diagnostics are a tool with variable and partially characterized performance, not a general solution to health inequity. Deploying them requires knowing where they perform reliably and where they don’t. It requires disclosure of performance limitations to clinicians and, eventually, patients. It requires ongoing surveillance that can detect performance problems in specific deployment populations before they translate into accumulated patient harm.

The democratization narrative is not false — it describes a genuine potential. It is also not self-fulfilling. Realizing that potential requires active effort to ensure that the populations who most need access to expert-level diagnosis are not systematically excluded from the populations these tools reliably serve.

That’s a harder project than building a good model. It is also, arguably, the more important one.

The Procurement Barrier

Health centers serving underserved communities face a procurement barrier that is rarely discussed in the AI health equity literature: the vendor contracting and implementation complexity required to deploy even a single AI diagnostic tool is substantial, and it disproportionately burdens the institutions least equipped to manage it.

Deploying an AI radiology tool requires IT integration with the PACS (Picture Archiving and Communication System), EMR integration for alert delivery, staff training, performance monitoring infrastructure, and ongoing vendor management for updates and performance surveillance. Each of these steps requires time, expertise, and budget that federally qualified health centers and rural critical access hospitals structurally lack. The vendor companies deploying AI tools have not, in general, designed their implementation processes around the constraints of under-resourced institutions.

Several academic medical centers have attempted to bridge this gap by creating “AI deployment consortia” — shared infrastructure services that smaller hospitals can access without building their own AI deployment capacity. The Mayo Clinic’s AI deployment licensing program and Partners HealthCare’s AI at Scale initiative are examples. The model is promising but reaches only a fraction of the institutions that would benefit.

The Regulatory Lever

There is a regulatory lever that could accelerate equitable deployment and is largely unused: tying FDA clearance or post-market surveillance obligations for AI diagnostic tools to demonstrated performance across diverse patient populations. Under current FDA frameworks, a device cleared based on a predominantly white, academic-center validation cohort can be deployed to any patient population without additional evidence of performance in that population.

Requiring prospective performance validation in demographically diverse cohorts before clearance — or requiring post-market surveillance studies in underrepresented populations as a condition of clearance — would shift the incentive structure for vendors. Tools with known performance gaps in specific populations would face regulatory pressure to address those gaps, not just market pressure from institutions that may not have the analytical capacity to detect them.

The FDA’s Digital Health Center of Excellence has signaled openness to this kind of requirement in its 2025 pre-submission guidance for AI medical devices. Whether that openness translates into binding requirements in cleared products will be determined in the next several regulatory cycles. The patients who would benefit most from getting this right are not the ones with the most powerful lobbying presence in the regulatory process — which is exactly why the oversight structure needs to be designed to protect them anyway.

The Vendor Responsibility Question

The gap between what AI vendors know about their tools’ limitations and what they disclose to deploying hospitals is a structural ethics problem that the field hasn’t adequately confronted. A company that has internal performance data showing their dermatology AI performs 15 percent worse in darker skin tones, but markets the tool without disclosing this, is not just creating legal exposure — it is making a choice to prioritize commercial deployment over patient safety in the populations most likely to lack alternative access to specialist care.

The argument that vendors can’t be held responsible for performance gaps that hospitals should detect through their own validation is technically defensible but practically bankrupt. Most community hospitals don’t have the staff, the labeled data, or the statistical expertise to run meaningful subgroup performance analyses on AI tools they’re considering procuring. They are depending on vendor disclosure to understand what they’re buying. When vendors don’t disclose, the harm lands on patients — specifically, on the patients in the demographic groups least represented in the vendor’s validation data, who are disproportionately patients in the community hospitals and health centers that can least afford the consequences.

This is a regulatory compliance problem, an ethics problem, and eventually a litigation problem. The sector will address it through one of those three vectors. The most efficient path is through regulatory compliance before the litigation makes it expensive and the ethics failure makes it a scandal.

Get the best of Think Different in your inbox

One email a month: new articles, reviews and the upcoming live webinar + free recording. No spam, unsubscribe anytime.