When the FDA’s Center for Drug Evaluation and Research began receiving New Drug Applications that contained AI-generated data in their supporting packages — molecular property predictions from ML models, patient selection algorithms for clinical trials, AI-assisted biomarker identification — the agency faced a problem that has no clean historical precedent.

Regulatory science, as practiced at FDA since the 1962 Kefauver-Harris amendments that formalized the modern approval process, is built around the concept of evidentiary reproducibility. A sponsor submits data. The FDA’s reviewers assess whether the data meets the evidentiary standards for safety and efficacy. The key assumption is that the methods used to generate the data can be independently validated and, in principle, reproduced. Statistical methods can be audited. Analytical chemistry can be verified. Clinical endpoints can be evaluated against pre-specified definitions.

Machine learning models — trained on proprietary datasets, with parameters numbering in the millions, producing predictions through processes that cannot always be articulated as interpretable decision rules — challenged several of these assumptions simultaneously. The FDA had to develop a theory of what it means to validate, trust, and approve data generated by a black box.

The Agency’s Response

The FDA’s response has been more methodical and substantive than the technology press (which prefers narratives about regulatory obstruction versus innovative technology) typically acknowledges. The agency established an Artificial Intelligence Center of Excellence within CDER in 2023, tasked with developing frameworks for evaluating AI contributions to drug development across the full spectrum from drug design to clinical trial conduct.

The first substantive guidance, published in draft form in 2024 and finalized in early 2026, addressed AI in drug development support — the use of ML tools in preclinical research, computational chemistry, and clinical trial design. The key regulatory principles it articulated were: transparency (sponsors must describe AI methods in sufficient detail that reviewers can assess their appropriateness), validation (AI tools used in support of regulatory submissions must be validated for their intended use with documented performance characterization), and risk-proportionality (the level of evidence required scales with the influence of the AI output on the clinical package).

This last principle is important. An AI tool used to predict metabolic stability in early-stage lead optimization, where the output informs but doesn’t determine compound selection, requires less formal validation documentation than an AI-derived patient selection biomarker used to define the primary efficacy population in a pivotal Phase III trial. The FDA is not demanding the same evidence for every AI application — it is trying to calibrate its scrutiny to the clinical consequence.

The Biomarker Validation Challenge

The hardest regulatory problem in AI drug development is biomarker validation. When a company uses an AI-derived multi-variable biomarker to select patients for a clinical trial, the FDA’s standard analytical validation framework (precision, accuracy, sensitivity, specificity — the classic metrics for a single-measurement biomarker) doesn’t translate cleanly. A multi-variable biomarker that combines genetic features, protein expression levels, imaging data, and clinical variables into a composite score is a computational object that requires a different validation framework.

The FDA’s Biomarker Qualification Program, which provides regulatory qualification for biomarkers used across development programs, has been navigating this challenge. The qualification of a complex multi-variable biomarker requires demonstrating its predictive validity across independent patient cohorts, characterizing its performance across subgroups (particularly demographic subgroups), and establishing analytical reproducibility — that the same patient sample analyzed in different laboratories produces the same result.

For AI biomarkers, analytical reproducibility includes reproducibility of the model itself across computing environments, which requires software validation practices imported from FDA’s guidance on software as a medical device. The intersection of drug development regulatory science and software device regulatory science is genuinely new territory, and the agency is developing it in real time with industry input.

The Trial Integrity Question

A less-discussed regulatory concern about AI in clinical trials is trial integrity. Clinical trials are adversarial environments in certain respects — sponsors have financial incentives to select patients and endpoints that maximize the probability of a positive result. Regulatory science has spent decades developing controls against bias: randomization, blinding, pre-specified analysis plans, independent data monitoring committees, and the general principle that the statistical analysis plan cannot be modified based on knowledge of outcomes.

AI tools introduce new vectors for potential bias in trial design and conduct. A machine learning model used to select patients for a trial can be, in principle, optimized to find the patient population where a drug appears most effective — a form of retrospective population selection that inflates the apparent treatment effect. The FDA is acutely aware of this risk and has issued guidance requiring that AI-derived patient selection criteria be specified and locked before any unblinded efficacy data is analyzed.

More subtle is the risk of data-driven endpoint selection. If a sponsor has access to a rich dataset of patient measurements and uses ML to identify which outcome measure is most strongly predicted by treatment in that dataset, then uses that outcome as the primary endpoint in a subsequent trial, they have effectively “mined” an endpoint from exploratory data — a practice that inflates false positive rates even when done without conscious intent to deceive. The FDA guidance requires that primary endpoints be defined on the basis of clinical meaningfulness and prior evidence, not selected based on AI-assisted analysis of current trial data.

International Regulatory Coordination

The FDA is not acting in isolation. The International Council for Harmonisation (ICH) has a working group specifically on AI in pharmaceutical development, with representatives from the FDA, EMA (European Medicines Agency), PMDA (Japan), Health Canada, and other agencies. The working group is developing international technical guidelines that would establish common standards for AI method documentation and validation across regulatory jurisdictions.

This matters because drug development is global and clinical data generated in one jurisdiction must often be accepted by regulatory agencies in others. If the FDA requires specific AI validation documentation that the EMA doesn’t require, sponsors face duplicative work and potential inconsistencies. The ICH process is slow — it typically takes 3-5 years from concept to finalized guideline — but it is the appropriate mechanism for achieving the regulatory harmonization that global drug development requires.

The EMA has been slightly more aggressive than the FDA in requiring AI transparency in submissions, reflecting the European regulatory culture’s stronger emphasis on explainability. The PMDA has been more conservative, reflecting Japan’s historically cautious approach to novel methodologies. These differences in regulatory philosophy are working themselves out in international guideline development rather than in conflicting country-level requirements, which is the optimal outcome.

What the Drug Approval Record Shows

As of mid-2026, no drug has been approved where the FDA explicitly credited AI as the primary discovery mechanism. Several approved drugs were discovered with AI assistance: FDA approval letters for oncology drugs in 2024-2025 reference computational structural biology and AI-assisted biomarker development in their summary review documents, acknowledging the methodology without specifically endorsing AI as a category.

The first truly pivotal regulatory test will come when an AI-first company files for approval of a drug where the AI contribution was essential and central — not just one tool among many, but the mechanism that found the target or designed the molecule. Insilico’s INS018_055 is the most likely candidate, if its Phase III succeeds. That trial’s regulatory review will be a reference event for how the FDA applies its developing AI framework to a real submission.

The agency is preparing for that moment, not reacting to it. Whether the preparation is sufficient is a judgment that will require watching the first few approvals and seeing how the framework holds under the stress of real submissions. The FDA has consistently surprised its critics with its ability to adapt to novel technologies — gene therapy, recombinant biologics, personalized genomic medicine — when the scientific evidence is clear. The institutional question is whether AI drug development will generate sufficiently clear evidence, and on what timeline.

The Staffing Challenge Inside FDA

Less discussed than the external regulatory framework is the internal challenge the FDA faces in building the expertise to evaluate AI drug development submissions. The agency’s reviewers are scientists with deep domain expertise in chemistry, pharmacology, clinical medicine, and biostatistics. AI drug development submissions require understanding not just the clinical data but the computational methods that generated or filtered it: the architecture of generative models, the validation of molecular property predictors, the statistical behavior of ML-derived biomarkers.

FDA has been actively recruiting computational scientists and data scientists into its review divisions, and the AI Center of Excellence is providing training and consulting to review teams across centers. But building genuine AI expertise in a federal workforce, against competition from industry salaries, is a structural challenge that training programs alone cannot solve. The agency depends substantially on external scientific advisory committees — which can draw on academic and industry expertise — to supplement internal reviewer competencies. Several recent AI-related advisory committee meetings have included explicit AI methodology sessions that would have been unusual five years ago.

This expertise gap creates risk not of arbitrary regulatory obstruction but of regulatory decisions made with insufficient technical depth on the AI components of complex submissions. A reviewer who doesn’t fully understand the behavior of a generative molecular design model may not know what questions to ask about its validation, which means they can’t identify the submission’s weakest points. The FDA’s response — hiring, training, advisory committee composition, and external collaborations with BARDA and NIH — is appropriate and ongoing. The gap is closing. It’s not yet closed.

The Anticipation of Adverse Events

The most important regulatory question in the medium term is what happens when an AI-designed drug causes a serious unexpected adverse event in a large post-approval population. The history of drug safety is full of harms that didn’t appear in clinical trials — Vioxx’s cardiovascular risk, which became apparent only in large post-marketing populations; thalidomide’s teratogenicity, famously — and post-market surveillance exists specifically to catch harms that pre-approval trials miss.

If an AI-designed drug causes a serious unexpected harm that is linked to a specific feature of the AI design process — a binding selectivity issue that the model didn’t predict, an off-target interaction that the ADMET models rated as low risk — the regulatory response will set precedents for AI drug development accountability. The FDA’s existing adverse event reporting framework is designed for any approved drug; it doesn’t have specific mechanisms for tracing harm back to computational design choices. Building that traceability — so that a post-market signal can be investigated by examining the AI system’s original design decisions and predictions — is a forward-looking regulatory challenge that the agency has begun addressing in its validation guidance but hasn’t fully resolved.

Get the best of Think Different in your inbox

One email a month: new articles, reviews and the upcoming live webinar + free recording. No spam, unsubscribe anytime.