Uncategorized

AI Sepsis Alert Tool Built at Duke Proves It Works Just as Well at Other Hospitals

A multisite validation study found Duke Health's Sepsis Watch AI model, credited with a 27% drop in sepsis deaths at Duke, performed just as strongly after being deployed at a different community hospital system.

AI Sepsis Alert Tool Built at Duke Proves It Works Just as Well at Other Hospitals

An artificial intelligence model developed at Duke Health to flag sepsis risk in emergency department patients has held up when transplanted into a completely different hospital system, according to a multisite external validation study published in npj Digital Medicine. The finding matters because AI clinical tools frequently perform far worse once they leave the hospital where they were built and trained.

The tool, called Sepsis Watch, was originally deployed at Duke Health starting in 2018 and has been associated with a 27% reduction in sepsis deaths there, according to prior reporting on the program. The new study set out to test whether that performance was a Duke-specific artifact or something that would generalize.

Testing the model somewhere new

Researchers evaluated Sepsis Watch’s portability by validating it in a community healthcare setting at Summa Health, analyzing 205,005 emergency department encounters involving 101,584 unique patients between 2020 and 2021, according to the published study and its preprint version on medRxiv. That is a substantially different patient population and care setting than Duke’s original academic medical center environment.

The model’s area under the receiver operating characteristic curve, or AUROC, a standard measure of how well a predictive model distinguishes patients who will develop sepsis from those who won’t, ranged from 0.906 to 0.960 across the sites tested — considered strong performance, and importantly, one that held steady rather than degrading sharply outside Duke.

Why generalization is the real bottleneck

The finding speaks to a well-documented weakness in clinical AI: models trained on one hospital’s patient mix, documentation habits, and lab equipment often lose accuracy elsewhere. The most cited cautionary example is the Epic Sepsis Model, a widely deployed proprietary tool used across hundreds of U.S. emergency departments that has been reported to experience performance degradation over time and across sites — a contrast researchers pointed to when framing why Sepsis Watch’s multisite consistency was noteworthy.

How widespread this kind of tool has become

Predictive AI embedded in electronic health records is now used by 71% of non-federal acute care hospitals in the U.S., up from 66% in 2023, according to federal data cited in recent nursing and health IT research. Separately, the FDA has authorized an AI biomarker tool called Sepsis ImmunoScore for identifying patients at elevated risk, and a 2025 meta-analysis spanning 52 studies found machine learning sepsis models consistently outperformed conventional risk-scoring tools, with AUC values ranging from 0.79 to 0.96 across the literature.

The limits proponents acknowledge

Even researchers behind Sepsis Watch’s validation are careful to note that strong AUROC numbers don’t automatically translate into better patient outcomes — that depends on whether clinicians act on the alerts in time and whether the alerts are specific enough to avoid contributing to alarm fatigue, a chronic problem in emergency departments already saturated with automated warnings. The study’s precision-recall figures, with AUPRC values between 0.177 and 0.252, were notably lower than the AUROC scores, a reminder that sepsis remains a relatively rare event even in large ED populations, which limits how precise any early-warning system can be in practice.

What comes next

With sepsis remaining one of the leading causes of in-hospital death and a frequent target of AI investment, the Sepsis Watch validation adds evidence that at least some sepsis-prediction models can travel between health systems without losing their edge — a prerequisite if hospitals are going to trust externally developed AI tools rather than building bespoke, locally trained models from scratch. Researchers involved in the work are expected to push for further validation across additional community and rural hospital settings, where sepsis mortality is often highest and specialist staffing is thinnest.

Why hospitals keep circling back to sepsis

Sepsis has long been considered one of the best test cases for clinical AI because it is common, frequently missed or caught late, and has a relatively well-defined bundle of interventions — antibiotics, fluids, closer monitoring — that improve outcomes if started early. That combination of high stakes and actionable next steps is precisely what has drawn so much AI investment to the condition, from academic efforts like Sepsis Watch to commercial and FDA-authorized tools like the Sepsis ImmunoScore biomarker and the proprietary Epic Sepsis Model built into one of the country’s most widely used electronic health record systems.

That crowded field also means hospitals evaluating a new sepsis AI tool today have real alternatives to compare against, and the Epic Sepsis Model’s documented performance decline over time has made many health system informatics teams more cautious about accepting vendor performance claims without independent, site-specific validation — exactly the kind of scrutiny the Sepsis Watch team invited by publishing its Summa Health results openly rather than relying on internal Duke data alone.