Uncategorized

A Swiss Hospital Built an AI That Watches Every Ward for Sepsis Every Six Hours

Lausanne University Hospital's AI system HERACLES re-screens every patient for sepsis every six hours. A study of nearly 100,000 patient stays found in-hospital mortality for flagged cases fell from 20.54% to 15.27%.

aidatanews

At Lausanne University Hospital in Switzerland, an algorithm named HERACLES quietly rechecks every patient in dozens of wards every six hours, sorting cases into confirmed sepsis, possible sepsis, or ruled out. A before-and-after study covering 97,559 patient stays across 63 wards running the system, compared against 25,851 stays in 126 control wards from 2020 to 2024, found that in-hospital mortality for HERACLES-flagged sepsis cases fell from 20.54% to 15.27%, while control wards trended upward over the same period, according to the study published in npj Digital Medicine and led by researchers including Jérémie Despraz, Jean Louis Raisaro and Sylvain Meylan as part of the CHUV Sepsis Group.

Sepsis Has Always Been a Race Against the Clock

Sepsis, the body’s often fatal overreaction to infection, kills an estimated 11 million people worldwide each year according to figures cited in prior sepsis research, and it remains notoriously hard to catch early because its warning signs, fever, rising heart rate, confusion, mimic dozens of less dangerous conditions. Hospitals have used manual screening tools like the Systemic Inflammatory Response Syndrome criteria and the quick Sequential Organ Failure Assessment score for years, but these depend on nurses and physicians remembering to check them consistently across a busy ward, and studies have repeatedly shown gaps in compliance. That inconsistency is exactly the gap machine learning models have spent the last decade trying to close.

How HERACLES Actually Works

Unlike a simple rules-based alert, HERACLES was trained on 1,043 expert-adjudicated cases, 472 confirmed sepsis, 341 possible sepsis and 230 non-sepsis, achieving an F1 score of 0.67 in validation, a measure that balances how often the model catches true cases against how often it wrongly flags healthy patients. The tool does not just fire a single alarm. It feeds a dynamic dashboard of quality-of-care indicators that clinical teams monitor continuously, re-classifying every patient stay across the hospital roughly four times a day rather than waiting for a nurse to manually trigger a screening protocol. That continuous re-evaluation is what the researchers describe as a “learning health system,” a model of care where the AI and the clinical pathway are designed and refined together rather than the algorithm being bolted onto an existing workflow after the fact.

The Numbers Behind the Headline

Beyond the in-hospital mortality drop, the study also tracked 90-day mortality, which fell from 32.99% to 26.11% for HERACLES-flagged patients in the intervention wards, according to the npj Digital Medicine paper. Sepsis coding, meaning how consistently clinicians documented sepsis diagnoses in patient records, also increased significantly in the AI-supported wards while remaining flat in control wards, suggesting the system may be catching cases that previously went formally undocumented as well as clinically undertreated.

Why Some Researchers Urge Caution

Before-and-after studies, even large ones spanning tens of thousands of patient stays, are widely considered weaker evidence than randomized controlled trials, because outcomes can shift over time due to unrelated factors like staffing changes, new antibiotic protocols, or even seasonal variation in the patient population, none of which the AI caused. Sepsis researchers publishing systematic reviews in outlets such as Frontiers in Digital Health have separately cautioned that many machine learning sepsis models look strong in retrospective validation but lose accuracy when deployed in real hospitals with different patient mixes, an effect sometimes called model drift. The Lausanne team itself frames its results as evidence the approach improves quality-of-care indicators within its own system, not as proof the specific F1 score of 0.67 will generalize cleanly to a hospital with a different patient population or a different electronic health record.

A Crowded, Competitive Field

HERACLES is not the only serious attempt at this problem. In the United States, an FDA-authorized tool called the Sepsis ImmunoScore has been evaluated in a prospective multi-center study across six hospitals for its ability to predict mortality in patients suspected of infection, according to research summarized on PubMed Central. Multiple academic groups are pursuing hospital-wide sepsis detection using natural language processing pulled directly from clinical notes rather than structured vital sign data alone. The field is converging on the same basic idea from several directions: continuous, automated re-screening beats relying on any single human to remember a checklist during a twelve-hour shift.

What Happens From Here

The Lausanne team’s next step, based on the trajectory of the published research, is expected to be extending the before-and-after design toward more rigorous prospective and potentially randomized comparisons, and testing whether the mortality improvements hold up as the algorithm is retrained on more recent patient data. For hospital administrators elsewhere, the appeal is straightforward: a tool that runs quietly in the background, checking patients whether or not an overworked staff member remembers to, without requiring new devices or wearables. Whether regulators in Europe and the United States move to formally authorize systems like HERACLES for wider adoption, the way the Sepsis ImmunoScore has already achieved FDA authorization, will likely determine how quickly this kind of ward-wide monitoring spreads beyond a handful of academic medical centers.