The most expensive failure in the pharmaceutical industry is a Phase III trial that fails. By the time a drug reaches Phase III, it has survived years of preclinical work and Phase I/II testing. The company has invested hundreds of millions of dollars. Patients are enrolled, expecting a chance at benefit. When the trial fails — as the majority of Phase III trials do — the cost is not just financial. It’s years of patient enrollment and the loss of a potential treatment.
The headline story of AI in drug development focuses on molecule generation and protein structure prediction. The less-publicized and arguably more impactful application is earlier in the clinical story: designing trials that are more likely to succeed, finding the patients most likely to respond, and identifying failure signals early enough that resources can be redirected rather than wasted.
The Patient Selection Problem
Most Phase III trial failures are not failures of the drug’s mechanism. The molecule did what it was supposed to do, in the patients for whom it worked. The problem is that those patients were diluted by a larger population of patients where the drug had no effect — patients who were enrolled in the trial because they had the same diagnosis, but whose disease was driven by a different biological mechanism than the one the drug targeted.
This is a well-known problem in oncology, where the lesson of the genomic era (since roughly 2010) was that “lung cancer” is not one disease but many diseases with overlapping radiological appearance and divergent molecular drivers. Giving an EGFR inhibitor to a patient with KRAS-mutant lung cancer doesn’t work, not because the drug is bad but because the drug targets the wrong mechanism for that patient.
The genomic revolution addressed this by identifying molecular biomarkers — gene mutations, copy number variations, protein expression levels — that predict response. But single-biomarker stratification only goes so far. Many diseases have complex, multi-factorial drivers that a single biomarker captures poorly. This is where machine learning adds genuine value: training on large datasets of patient characteristics, treatment histories, and outcomes to build multi-variable predictive models that identify responsive subpopulations with better precision than single biomarker tests.
Several clinical-stage programs have used ML-derived biomarkers to select patient populations in Phase II with striking results. Relay Therapeutics’ RLY-4008 (FGFR2 inhibitor for bile duct cancer) enrolled patients using a combination of FGFR2 molecular testing and an ML model predicting resistance pathways, and achieved a response rate that exceeded any previous single-agent targeted therapy in this indication. Merus NV used a computational approach to identify patients most likely to respond to their MCLA-158 bispecific, based on expression profiles that went beyond simple biomarker positivity. Early Phase II data was sufficiently strong to accelerate the Phase III design.
Trial Protocol Optimization
The other hidden contribution of AI to clinical trials is protocol optimization — the design of the trial itself. Trial protocol design involves dozens of decisions that collectively determine efficiency and statistical power: eligibility criteria, primary and secondary endpoints, dose selection, dosing schedule, patient monitoring requirements, trial duration. Historically, these decisions were made by clinical teams based on experience, regulatory guidance, and competitive precedent, with limited quantitative support.
AI tools are changing this in two ways. First, by mining historical trial databases (FDA submission records, ClinicalTrials.gov data, published literature) to identify protocol features associated with success or failure across comparable compounds. A team designing a Phase II trial for a novel JAK inhibitor can now query ML models trained on all previous JAK inhibitor trials to identify which endpoint choices, which patient populations, and which dosing designs had the best success rates — not based on intuition but on statistical patterns across hundreds of precedent programs.
Second, by using simulation to stress-test protocol assumptions before enrollment begins. Clinical trial simulation has been part of statistical methodology for decades, but the models were limited by the quality of prior data inputs. ML-enhanced simulation can incorporate more complex prior data, including real-world electronic health record data from similar patient populations, to predict trial behavior under different protocol assumptions with greater fidelity. Several companies (Unlearn.ai most visibly) are specifically building on this for control arm optimization — using synthetic digital twins of patients, generated from historical data, to reduce the size of the placebo arm and thereby require fewer patients to reach statistical power.
The Dropout Prediction Problem
Clinical trial dropout — patients who enroll and then discontinue before completing the protocol — is a chronic, expensive, and underappreciated source of trial failure. A trial designed for 300 patients with 10 percent assumed dropout may be powered inadequately if actual dropout is 25 percent. Dropout analysis is often retrospective, after the trial has already been impacted.
Predictive models trained on electronic health records from prior trial participants can identify, at enrollment, which patients are at high risk for dropout based on demographic factors, disease characteristics, geographic distance from sites, prior medication adherence patterns, and other variables. Using this to inform site selection, protocol design (reducing visit burden for high-risk populations), and targeted retention interventions could improve trial completion rates and reduce the waste of failed trials due to statistical underpowering.
This is less headline-worthy than AI-designed molecules. It’s also more immediately deployable: it doesn’t require novel chemistry or regulatory innovation. The data exists in EMR systems. The models can be trained now. The implementation barrier is primarily organizational, not technical.
The Adaptive Trial Revolution
Adaptive trial designs — which use pre-specified rules to modify trial parameters based on interim results — have been advocated by statisticians for decades and resisted by regulators and pharmaceutical companies who found them operationally complex and analytically fraught. AI is beginning to change the feasibility calculation by automating the real-time data analysis required for adaptation decisions.
The FDA’s Complex Innovative Trial Design (CID) program, active since 2019, has progressively expanded the regulatory pathway for adaptive designs. In 2024-2025, several adaptive Phase II/III seamless designs received regulatory concurrence — trials that begin as Phase II studies and can expand directly into Phase III without a protocol amendment if interim results meet pre-specified criteria. These designs can shorten development timelines by 12-18 months for successful drugs.
The AI contribution here is less in the statistical methodology (which is mature) and more in the real-time data processing and quality monitoring that makes adaptation decisions reliable. Adaptive trials require rapid, clean interim datasets, and the historical challenge was that data cleaning and lock processes took months — making “interim analysis” a near-oxymoron in practice. AI-assisted data monitoring and automated query resolution has cut data lock times at some sites from weeks to days.
The Regulatory Dimension
All of this is occurring against the backdrop of regulatory agencies — FDA, EMA, PMDA — developing frameworks for validating and trusting AI contributions to drug development. The FDA’s 2024 guidance on AI in drug development explicitly addresses using AI in trial design and patient selection, requiring sponsors to describe and validate their AI methods with the same rigor applied to other analytical tools.
This is appropriate and important. An AI-selected patient population can only support labeling claims for that population if the selection method has been validated. A Phase III trial designed using ML protocol optimization tools must document and justify those tools’ contribution. The regulatory agencies are not blocking AI applications in clinical development; they are asking for the same evidentiary standards that apply to any other development tool.
The companies that will benefit most from AI in clinical development are those building the documented, validated pipelines that regulatory agencies can audit and approve — not those treating AI tools as black boxes whose contribution is hard to trace. The regulatory pathway rewards transparency, and transparency in AI methods requires investment in interpretability and documentation that many organizations have treated as optional. In clinical development, it isn’t.
The Real-World Evidence Integration
One frontier that is still nascent but potentially transformative is the integration of real-world evidence (RWE) — data from electronic health records, insurance claims, wearable devices, and patient registries — into trial design and analysis. Traditional trials generate their own data in controlled environments. RWE provides evidence about how drugs perform in diverse, uncontrolled real-world patient populations.
AI is essential to using RWE effectively because the data is messy: variable documentation quality, inconsistent coding practices, missing data, and confounding that controlled trial design prevents but observational data cannot. ML methods for causal inference from observational data — inverse probability weighting, propensity score matching, double machine learning — are being applied to RWE to generate synthetic control arms, validate trial enrollment criteria in real-world patient populations, and estimate how trial results might generalize to patients who don’t look like the typical trial participant.
The FDA’s Real-World Evidence Program, active since 2017 and substantially expanded through 2025, has approved several RWE-supported labeling expansions. The evidence bar for RWE is still significantly higher than for randomized controlled trials, and appropriately so. But the direction of regulatory travel is toward accepting well-executed RWE as supportive evidence for specific questions — particularly post-approval, in populations underrepresented in registration trials.
The Cost Question
The clinical trials that AI is making more efficient are still extraordinarily expensive. Phase III trials typically cost $200-600 million, and the cost has been rising at roughly 8 percent annually for the past decade despite efficiency improvements in some components. AI can reduce specific costs — patient recruitment time, data cleaning labor, protocol deviation rates — but it doesn’t address the fundamental cost driver: the size of the patient population required to demonstrate statistical significance on clinical endpoints that medicine considers meaningful.
That cost driver is a function of biology (how heterogeneous the patient population is, how strong the treatment effect is, how variable the outcome measure is) and regulatory standards (how much evidence of benefit FDA requires before approving a drug). AI can improve patient selection to reduce heterogeneity, and it can help identify biomarkers that improve treatment effect estimates. But it cannot lower regulatory standards or fundamentally change the biology of disease heterogeneity. The trials will remain expensive. The question is whether they become more expensive more slowly, and whether they succeed more often — which is the actual prize.
One email a month: new articles, reviews and the upcoming live webinar + free recording. No spam, unsubscribe anytime.