A Cleveland Clinic-affiliated research team has published a foundation model in Nature Communications on August 6, 2026, that decodes overnight sleep-study data into detailed risk categories for death, heart disease, and cognitive decline, outperforming the decades-old apnea-hypopnea index that doctors currently rely on to grade sleep apnea severity. The model, described in the paper “A foundation model for sleep-based risk stratification and clinical outcomes,” was trained on 9,608 full-night polysomnography studies from 9,297 patients drawn from the Cleveland Clinic’s Sleep Signals, Testing, and Reports Linked to Patient Traits registry, known as STARLIT.
Why the Old Scoring System Falls Short
For decades, clinicians have graded sleep apnea severity almost entirely through the apnea-hypopnea index, a single summary number counting how many times per hour a patient’s breathing stops or nearly stops during sleep. That number is simple to calculate but, according to the new research, captures only a sliver of the physiological information contained in a full night of sleep data, missing subtler patterns in brain activity, oxygen fluctuation, and sleep architecture that carry independent health risk. The Cleveland Clinic team’s central finding underscores the gap starkly: patients in the model’s highest-risk group faced more than double the mortality risk of those in the lowest-risk group, while the traditional apnea-hypopnea index, applied to the same patients, “failed to detect any association with mortality” at all.
How the Model Was Built and Tested
Researchers linked the polysomnography data to electronic medical records spanning an average follow-up of 14.5 years, letting them track which sleep-derived risk categories actually predicted which real-world health outcomes over time rather than relying on short-term proxies. The team, led by researcher Bilal E., then validated the model against an independent dataset, the long-running Sleep Heart Health Study, which used different polysomnography equipment and protocols than the Cleveland Clinic’s own STARLIT registry. That cross-dataset validation matters because AI models trained on one hospital’s specific equipment and patient population frequently fail to generalize when applied elsewhere, a common criticism of single-site medical AI research.
Five Risk Groups With Distinct Trajectories
Rather than producing one severity score, the foundation model sorted patients into five distinct risk groups, each showing a different combination of outcomes across mortality, major adverse cardiovascular events, atrial fibrillation, cognitive impairment, and even epilepsy risk. Critically, the researchers reported that these AI-derived groupings held up statistically “even after adjusting for demographics, comorbidities, and the AHI itself,” meaning the model was picking up genuinely new risk information rather than simply repackaging factors doctors already track through age, weight, or existing diagnoses.
Part of a Broader Rethink of What Sleep Data Reveals
The Cleveland Clinic findings arrive alongside other 2026 research pushing sleep science well beyond apnea counting. A Stanford Medicine study published earlier in the year found that a single night of sleep data could help predict risk across more than 130 distinct health conditions, including cancer, heart disease, and dementia, using patterns invisible to conventional scoring. Consumer smartwatch makers have moved in parallel: both Samsung and Apple received FDA clearance in 2026 for smartwatch-based sleep apnea detection, pushing screening capability out of overnight sleep labs and into everyday wearables, even as the underlying clinical interpretation tools researchers are now questioning remain largely unchanged.
Two Views on How Fast This Should Reach Patients
Sleep medicine researchers championing the foundation model argue it exposes a real, previously invisible blind spot in how millions of sleep studies get interpreted every year, since a normal or mild AHI reading today may be reassuring patients and doctors incorrectly about mortality risk the AI model can already detect. More cautious voices in the field, including the study’s own authors, note that prospective clinical trials and further external validation are required before any foundation model like this could be used to actually change a patient’s diagnosis or treatment plan, and that additional work is needed on self-supervised and disease-targeted training methods to reduce the model’s current dependence on labels defined by human sleep technicians.
What Happens Next
The Cleveland Clinic team has flagged prospective trials as the next necessary step before any version of this foundation model could move from research finding to bedside tool, a process that typically takes years given the need to track real patient outcomes forward rather than retrospectively, as this study did. In the meantime, the results add pressure on sleep medicine as a field to reconsider whether a single summary number, unchanged in its basic form since the apnea-hypopnea index was first adopted decades ago, is still the right way to communicate risk to the millions of patients who undergo an overnight sleep study each year.