Epic Systems — which provides electronic health records to roughly 35 percent of U.S. hospitals and a growing share of European and Asia-Pacific institutions — added a sepsis prediction algorithm to its platform in 2017. The tool, called the Sepsis Prediction Model (known informally as the Epic Deterioration Index), generated alerts for nurses and physicians when a patient’s vital signs, labs, and other EMR data matched patterns associated with sepsis onset. By 2022, hundreds of hospitals were running it. By 2025, millions of patients had been flagged by it.

In 2021, a team of researchers at the University of Michigan published a study in JAMA Internal Medicine that evaluated the Epic sepsis model’s performance at their institution. The findings were not reassuring: the model missed approximately 67 percent of sepsis cases while simultaneously generating false alarms for 18 percent of non-sepsis patients. A pair of independent evaluations at other institutions found broadly similar results. Epic’s model — deployed at scale, trusted by hospitals on the basis of Epic’s reputation and the face validity of the concept — was performing materially worse than clinicians had been told to expect.

This is a story about what happens when clinical AI scales faster than the evidence base that should support it.

Why the Sepsis Model Underperformed

The Epic sepsis model was trained on Epic’s proprietary dataset from specific institutions and validated on that dataset. When deployed to different hospitals with different patient populations, different documentation patterns, different clinical workflows, and different baseline sepsis rates, its performance degraded substantially. The problem is not unique to Epic — it is the distribution shift problem that affects all clinical ML models deployed outside their development context.

Sepsis is also a clinical syndrome defined by a combination of infection and dysregulated host response — a definition (Sepsis-3, adopted in 2016) that captures a spectrum of severity rather than a single biological entity. Training a model on “sepsis” labels in one institution means training on that institution’s clinicians’ application of that definition to their patients. When applied to a different institution’s patients under different clinicians’ interpretations of the same definition, the model is predicting its training data, not the underlying biology.

The practical consequence is a high false alarm rate that clinical teams learn to ignore. In hospitals where the Epic sepsis alert fires many times per day and is wrong most of the time, the alert has been documented in multiple studies to generate alarm fatigue — clinicians devalue the signal because its positive predictive value is too low to justify the cognitive cost of acting on each alert. An alert that nobody trusts is worse than no alert, because it consumes attention that might otherwise be applied to actual clinical assessment.

The Alert Fatigue Epidemic

Alarm fatigue is not a new problem in medicine — it predates AI by decades. Cardiac monitors in ICUs, pulse oximetry alarms, infusion pump alerts — all of these generate high volumes of alerts the majority of which do not require action, and clinicians have developed organizational strategies to manage the noise (often including silencing alarms or ignoring them by default, sometimes with fatal consequences).

AI clinical decision support adds a new layer of alerts on top of this existing problem. Most hospitals that implement AI deterioration prediction, sepsis alerts, or readmission risk scores implement them without comprehensively auditing their existing alert burden or evaluating the incremental attention cost of new alert types. The result is additive alarm load in environments where alarm load was already contributing to errors and burnout.

A 2024 systematic review in Critical Care Medicine evaluated studies of AI deterioration prediction in adult inpatient settings and found that of 28 studies reporting on clinical outcomes (not just predictive accuracy), only 9 found statistically significant improvements in outcomes attributable to the AI tool. The majority of studies reported on AUC (area under the receiver operating characteristic curve) — a measure of discrimination — without evaluating whether the tool actually improved patient care.

This gap between predictive accuracy and patient outcome is fundamental. A model that correctly identifies patients at risk of deterioration 12 hours before it happens is only clinically useful if the hospital can do something meaningful with 12 hours of lead time. If the nurse who receives the alert is already at capacity, if the intervention options are limited, if the physician workflow doesn’t have a mechanism to act on an AI alert without disrupting existing clinical priorities — the prediction goes unused. The clinical utility of a predictive tool depends entirely on the clinical system it is embedded in, not just on the model’s statistical performance.

The Mortality Prediction Problem

End-of-life decision-making in the ICU is one of the most emotionally and ethically complex areas of medicine. Prognostication — estimating the probability that a patient will survive an ICU stay — is central to conversations about goals of care, withdrawal of life support, and patient and family decision-making.

AI mortality prediction models (like those derived from APACHE, SOFA, and more recently ML-enhanced tools like TREWS and AKI predictors) can predict in-hospital mortality with reasonable statistical accuracy at the population level. The question that is not asked often enough is how those predictions should be used in individual patient conversations.

A model that predicts an 80 percent probability of in-hospital mortality for a patient is expressing statistical uncertainty. Among 100 patients with that predicted probability, roughly 80 die and 20 survive. The 20 survivors cannot be identified in advance. When a physician presents this probability to a patient’s family, the family is being told something statistically accurate and prognostically useful. They are also being told something with a 20 percent chance of being wrong for their specific family member in a way that matters enormously.

There is documented evidence that presenting algorithmic mortality predictions to clinical teams systematically affects care decisions in ways that disadvantage certain patient groups. A 2024 investigation of AI-generated mortality scores in a large teaching hospital found that patients whose AI-generated mortality risk was above a threshold (not publicly visible but computationally accessible to case managers) were significantly more likely to have goals-of-care conversations initiated in the first 48 hours of ICU admission, after controlling for documented clinical severity. The algorithm was influencing care patterns in ways that were not transparent to patients or families.

What Should Be Different

The deployment of AI clinical decision support in critical care is largely ungoverned at the level of clinical practice. Unlike a pharmaceutical drug (which requires FDA approval for each indication and rigorous post-market surveillance) or a diagnostic imaging device (which requires 510(k) clearance and premarket notification), clinical decision support software deployed through an EMR has historically operated under a regulatory carve-out that exempts it from FDA oversight when it merely assists clinicians without replacing clinical judgment.

The FDA’s 2024 Clinical Decision Support guidance narrowed this exemption somewhat, requiring that high-risk decision support tools — including sepsis prediction and deterioration algorithms that influence critical care decisions — undergo premarket review. The practical effect of this change is still unfolding, as existing tools grandfathered before the guidance are not automatically subject to review.

The deeper question is not regulatory but institutional: before a health system deploys an AI clinical decision support tool to its ICU, what validation should it require? What performance threshold? What evidence of generalizability to its specific patient population? Most hospitals don’t have the expertise to answer these questions, and the EMR vendor selling them the tool has limited incentive to make the answer inconvenient.

The ICU is where patients are most vulnerable and the consequences of getting care wrong are most severe. It is exactly the setting that deserves the most rigorous evidence standards for AI deployment. The deployment reality, so far, has not reflected that priority.

The Positive Cases

To avoid painting an entirely negative picture: there are ICU AI applications where the evidence base is stronger and the benefit more clearly demonstrated. AKI (acute kidney injury) prediction is probably the most validated application. A 2019 study in Nature by DeepMind trained a model on US Veterans Administration patient data and demonstrated AKI prediction up to 48 hours before clinical diagnosis, with potential to prevent irreversible kidney damage. A subsequent deployment study in a UK NHS trust found that early warning alerts generated by the algorithm, integrated with clinical decision support, were associated with earlier nephrology consultation and improved fluid management practices.

AI-assisted ventilator management — algorithms that recommend ventilator parameter adjustments for mechanically ventilated ICU patients — has been validated in several academic centers. The complexity of optimizing ventilator settings (pressure, rate, FiO2, PEEP) for individual patients with lung disease is genuinely high, and clinical judgment about optimal settings varies substantially across intensivists. AI systems trained on large ventilator datasets and patient outcomes can identify parameter combinations that are systematically associated with better outcomes. One randomized trial in France (2024) found that ventilator management guided by an AI recommendation system reduced ICU length of stay by 1.4 days on average in patients with ARDS — a meaningful clinical endpoint.

The distinction between these positive examples and the problematic applications is not algorithm quality per se. It’s specificity of the clinical problem, quality of the training data, and the existence of randomized controlled evidence demonstrating patient benefit rather than just predictive accuracy. Building on those positive cases — rather than generalizing “ICU AI” from the well-validated applications to the poorly-validated ones — is the discipline that clinical AI governance needs to enforce.

The Staffing Reality

Any honest account of AI in critical care has to acknowledge the context driving its adoption: ICU staffing is in crisis in most Western health systems. The United States is short approximately 9,000 intensivists (critical care physicians) relative to projected need. Nursing vacancy rates in ICUs averaged 18 percent nationally in 2025. The appeal of AI tools that can extend the reach of limited clinical staff — flagging deteriorating patients so the one available intensivist can prioritize, automating documentation so nurses can spend more time on direct patient care — is not irrational.

But the staffing crisis doesn’t change the evidentiary standard for what AI tools should be deployed in critical care. It creates pressure to deploy tools before that standard is met. Resisting that pressure — insisting on randomized controlled evidence for high-risk applications, limiting deployment to tools with characterized performance in comparable patient populations — requires institutional leadership that is difficult to maintain when the alternative is simply having fewer nurses and physicians trying to care for the same number of patients. The technology will continue to advance into this vacuum. The governance frameworks need to be built before the vacuum becomes the deployment rationale.

Get the best of Think Different in your inbox

One email a month: new articles, reviews and the upcoming live webinar + free recording. No spam, unsubscribe anytime.