A sleep study cannot currently tell you whether or when you are going to die. But a new peer-reviewed study found something that may ultimately be more useful: the physiological signals recorded during a standard overnight sleep study contained information about future mortality and disease risk that conventional sleep-apnea severity scores failed to capture.
Researchers analyzing thousands of Cleveland Clinic sleep studies used an AI foundation model to identify five distinct physiological groups. After adjustment for age, gender, body mass index, existing health conditions and the traditional apnea-hypopnea index, or AHI, patients in the highest-risk group had approximately 2.38 times the mortality hazard of patients in the lowest-risk group.
The result was not limited to the Cleveland Clinic dataset. When researchers tested their approach against the independent Sleep Heart Health Study, higher-risk groups again had greater mortality, although the effect was smaller: the highest group had an adjusted mortality hazard ratio of about 1.58 compared with the lowest.
That is significant evidence that a night’s sleep contains health information medicine is not fully using.
It is not evidence that an AI can examine your Apple Watch data and announce that you are going to die.
Those are very different claims.
What did the new AI sleep study actually find?
The study, “A foundation model for sleep-based risk stratification and clinical outcomes,” was published in Nature Communications on August 3, 2026.
Researchers began with 10,000 overnight polysomnography studies from Cleveland Clinic’s STARLIT registry. After quality-control processing, the main clustering analysis included 9,608 sleep studies from 9,297 patients.
A polysomnogram, or PSG, is far more comprehensive than simply measuring how long someone sleeps. Clinical sleep studies can simultaneously record brain waves, eye movements, muscle activity, heart electrical activity, breathing, respiratory effort and blood oxygen. The Cleveland Clinic recordings included EEG, EOG, EMG, ECG, oxygen saturation, nasal pressure, airflow, chest and abdominal movement, snoring and carbon-dioxide measurements.
Researchers trained a transformer-based foundation model to recognize three components of sleep physiology:
- sleep stages;
- respiratory events such as apneas and hypopneas; and
- oxygen desaturation events.
The important part is what happened next.
The researchers used the model’s internal representations of sleep physiology—often called embeddings—to group patients according to similarities in their overnight physiological patterns.
The model ultimately produced five stable clusters.
Only afterward did researchers examine how those groups differed in mortality and disease outcomes. The clusters were then labeled risk groups RG1 through RG5 based on those observed outcomes.
That means this was not simply an algorithm given a database of dead and living patients and trained to guess who would die.
It first learned representations of the physiology occurring during sleep. Researchers then discovered that those physiological patterns separated people into groups with substantially different health trajectories.
How much higher was the risk of death?
Mortality increased progressively across the AI-derived groups.
In the fully adjusted analysis, using the lowest-risk group as the reference:
- RG2: hazard ratio 1.43
- RG3: hazard ratio 1.54
- RG4: hazard ratio 1.75
- RG5: hazard ratio 2.38
Those models adjusted for demographics, BMI, relevant comorbidities and AHI. The association therefore persisted even after accounting for conventional sleep-apnea severity.
The pattern matters almost as much as the largest number.
If the result were simply an artifact dividing one unusual subset from everybody else, the finding would be less persuasive. Instead, mortality generally rose as patients moved through the AI-derived risk groups.
That is consistent with the model identifying a physiological risk gradient rather than merely an arbitrary category.
But the 2.38 hazard ratio is frequently misunderstood.
It does not mean someone in RG5 had a 238% probability of dying.
It does not mean 2.38 times as many people necessarily died.
And it certainly does not mean the algorithm could tell an individual patient, “You are going to die.”
A hazard ratio compares the rate at which an outcome occurs between groups over time, after whatever statistical adjustments were included in the model. It is a measure of relative risk over the observation period, not a personalized countdown clock.
The study did not predict deaths 14.5 years into the future
There is another easy way to overstate the study.
The Cleveland Clinic dataset had a mean total medical-record observation period of approximately 14.5 years. But that number includes health records from both before and after each patient’s sleep study.
Patients had a median 9.4 years of medical history before their polysomnogram and a median 4.4 years of follow-up afterward. In the mortality analysis itself, mean follow-up after the sleep study was approximately 4.9 years.
So it would be misleading to say:
“The AI looked at one night of sleep and predicted who would die 14.5 years later.”
That is not what the study demonstrated.
The evidence supports a narrower but still important conclusion: physiological information from a night’s sleep was associated with markedly different mortality risk during subsequent follow-up.
What did the AI see that the standard sleep-apnea score missed?
This may be the most important finding in the paper.
Sleep-apnea diagnosis and severity have traditionally relied heavily on the apnea-hypopnea index, or AHI.
AHI is essentially the number of apneas—periods when breathing stops—and hypopneas—periods of substantially reduced breathing—per hour of sleep.
It is useful. But it also compresses an enormously complicated biological process into a single number.
Consider what happens during a full-night polysomnogram.
Doctors can record:
- electrical activity in the brain;
- sleep-stage transitions;
- heart rhythms;
- muscle activity;
- oxygen levels;
- respiratory effort;
- airflow;
- awakenings and arousals;
- patterns of sleep fragmentation;
- interactions among all of those systems throughout the night.
AHI asks a narrower question:
How frequently did certain breathing events occur?
The new AI model was capable of looking across far more of the physiology simultaneously.
And the results suggest information is being lost when the entire night is reduced to conventional summary measures.
AHI severity did not predict mortality in this cohort
The researchers directly compared their AI-derived groups with conventional AHI categories.
The standard categories—normal, mild, moderate and severe sleep apnea—did not significantly separate mortality risk in the Cleveland Clinic cohort after adjustment for the same general covariates. The AI-derived risk groups did.
Even patients with severe AHI scores were distributed across several AI risk groups. Conversely, the highest-risk AI group contained patients from different AHI categories.
In other words:
Having more breathing interruptions did not automatically place someone in the AI model’s highest-risk physiological category.
That does not mean AHI is useless.
It means AHI and the AI model were measuring different things.
AHI remains clinically important for diagnosing and classifying sleep-disordered breathing. The study instead raises a different question: Is AHI too crude to serve as a general measure of the health consequences encoded in a night’s sleep?
There is already substantial interest within sleep medicine in answering that question. Researchers have investigated alternative measures such as hypoxic burden, heart-rate responses and sleep fragmentation because people with similar AHI scores can experience very different symptoms and health outcomes.
The new work suggests AI may be able to integrate several such dimensions simultaneously rather than forcing doctors to choose one summary metric.
The model was looking at more than breathing
Another important test was whether the AI had essentially invented a more complicated sleep-apnea score.
The evidence suggests it had not.
Researchers performed sensitivity tests in which they removed EEG brain-wave information or ECG cardiac information when generating the model’s physiological representations.
Doing so substantially changed which groups patients were assigned to.
That indicates the system was extracting information from multiple physiological systems—not merely counting respiratory interruptions differently.
The highest-risk group also showed increased incidence of several cardiovascular and neurological outcomes. In the study’s adjusted analyses, RG5 was associated with higher risks of major adverse cardiovascular events, heart failure, myocardial infarction, atrial fibrillation, cognitive impairment and epilepsy compared with RG1.
Those associations do not prove that disturbed sleep caused those diseases.
They suggest the physiology expressed during sleep may function as a kind of overnight stress test that exposes abnormalities occurring across the brain, heart, respiratory system and autonomic nervous system.
What sleep pattern predicts death?
There is no single “death pattern” identified by this study.
That is important.
Someone cannot look at their sleep duration, REM percentage or oxygen level and determine whether they belong to the high-risk group.
The five-group model was specifically valuable because ordinary PSG measurements could not reliably reproduce its more detailed classifications.
A simpler two-group separation was strongly associated with a measure the researchers call spectral sleep fragmentation, which quantifies rapid transitions between sleep stages. In the external Sleep Heart Health Study, people classified in the higher-risk group also tended to have markedly shorter total sleep times.
But neither observation establishes a rule such as:
“If you sleep fewer than X hours, your mortality risk doubles.”
The AI’s more detailed classifications depended on multidimensional physiological patterns.
Previous research does, however, point in a similar direction. A 2022 deep-learning study using polysomnography found that people whose sleep physiology appeared biologically “older” than their chronological age had higher mortality. Sleep fragmentation was one of the strongest features associated with that difference.
So the emerging pattern is not that scientists have discovered one fatal sleep behavior.
It is that sleep may reveal the accumulated condition of multiple physiological systems more clearly than conventional sleep metrics have been able to measure.
Did the result work in another population?
Yes, and this substantially strengthens the study.
Researchers applied the risk-stratification framework to the independent Sleep Heart Health Study, a large population-based research cohort.
The two datasets were meaningfully different. Cleveland Clinic patients were people undergoing clinical sleep testing, while the Sleep Heart Health Study was a community-based research cohort. The external recordings also contained fewer and lower-resolution channels than the Cleveland Clinic polysomnograms.
Despite those differences, mortality again increased among the higher-risk groups.
Compared with RG1:
- RG4 had a mortality hazard ratio of 1.31;
- RG5 had a mortality hazard ratio of 1.58.
The highest-risk group also had approximately 2.13 times the hazard of incident heart failure compared with the lowest-risk group in that validation cohort.
The smaller mortality effect in the external population is worth emphasizing.
The Cleveland Clinic result of 2.38 should not simply be treated as a universal multiplier that applies to everybody.
Replication of the direction of the association is encouraging. The reduced magnitude is also a reminder that effect sizes can change when a model moves into a different population.
Both facts matter.
Can your Apple Watch predict whether you’re going to die?
No—not from this research.
The study analyzed clinical polysomnography.
An Apple Watch does not record the same set of signals.
Apple’s sleep-apnea notification feature, for example, uses its accelerometer to detect wrist movements associated with breathing disturbances. It analyzes repeated measurements over a 30-day period to look for signs consistent with moderate to severe sleep apnea. Apple explicitly says the feature is not intended to replace medical diagnosis.
That is fundamentally different from a laboratory polysomnogram containing EEG brain activity, ECG cardiac activity, respiratory effort, airflow, muscle signals and multiple additional physiological channels.
Can an Oura Ring predict mortality from sleep?
There is no evidence from this study that it can.
Oura measures useful sleep and physiological signals, including movement, heart-related metrics and blood oxygen, but it does not collect the same full polysomnographic dataset used to develop this model.
Oura itself states that the Ring is not intended to diagnose, treat, cure, monitor or prevent medical conditions.
The same basic limitation applies to other consumer sleep trackers.
Their data may eventually contribute to disease-risk models. Researchers are actively studying that possibility.
But the August 2026 study does not establish that a Fitbit, WHOOP, Apple Watch, Oura Ring or sleep-tracking phone app can determine someone’s mortality risk.
This is not the first AI to find mortality information in sleep
Another misleading interpretation of the headlines would be that researchers have only now discovered that sleep physiology can predict mortality.
They have not.
In 2022, researchers reported that deep-learning estimates of biological age derived from polysomnography were associated with life expectancy and mortality.
And in January 2026, researchers published SleepFM in Nature Medicine, a much larger sleep foundation model trained on more than 585,000 hours of sleep recordings from approximately 65,000 participants.
From a single night’s polysomnography, SleepFM predicted a wide range of future outcomes, including all-cause mortality, dementia, myocardial infarction, heart failure, chronic kidney disease, stroke and atrial fibrillation. Its reported concordance index for all-cause mortality was 0.84.
So the new Cleveland Clinic/IBM study should not be described as proof that scientists suddenly discovered sleep can reveal mortality risk.
Its contribution is more specific.
It provides further evidence that foundation models can extract clinically meaningful physiological subtypes from routine sleep studies, and that those subtypes can outperform conventional AHI categories in distinguishing long-term health risk.
That distinction is less sensational.
It is also scientifically more interesting.
Does this mean the apnea-hypopnea index is obsolete?
No.
The new study does not establish that AI should replace AHI for diagnosing sleep apnea.
AHI answers an important clinical question: how frequently apnea and hypopnea events occur during sleep.
The AI model is aimed at a broader question:
What does the full pattern of someone’s overnight physiology say about their future health?
Those questions overlap, but they are not identical.
The potential future model of sleep medicine may therefore be additive rather than substitutive.
A patient could still receive an AHI for diagnosing and managing sleep-disordered breathing, while software simultaneously analyzes the full polysomnogram for patterns associated with cardiovascular, neurological or mortality risk.
That would turn the sleep study from primarily a diagnostic test for sleep disorders into something closer to a multisystem physiological assessment conducted while the patient sleeps.
The evidence is not yet strong enough to make that standard medical practice.
But the direction of the research is becoming increasingly difficult to dismiss.
Could doctors start getting an AI mortality score after a sleep study?
Possibly eventually.
Not yet.
The researchers themselves say additional replication and prospective clinical trials are needed before this type of system is deployed in healthcare.
That distinction matters because a model can identify statistical risk without proving that acting on the result improves anyone’s health.
Suppose a patient’s sleep study places them in a high-risk physiological group.
What should the doctor actually do?
Order a cardiology evaluation?
Perform neurological testing?
Repeat the sleep study?
Change medication?
Treat sleep apnea more aggressively?
Monitor the patient more frequently?
Nobody has yet demonstrated through prospective trials that using this AI classification to make those decisions improves outcomes.
Until that happens, calling it a validated clinical mortality test would be premature.
What are the study’s biggest limitations?
The finding is substantial, but there are several reasons not to treat it as a finished clinical technology.
The main Cleveland Clinic analysis was retrospective
Researchers analyzed existing sleep studies and subsequent medical records.
Retrospective data can identify strong associations, but it cannot prove that the physiological patterns detected by the AI caused the later diseases or deaths.
Treatment information was incomplete
The researchers did not have objective data showing how consistently patients used positive-airway-pressure therapy.
They performed sensitivity analyses excluding patients with PAP prescriptions and still found similar results, but incomplete treatment data remain a possible source of confounding. Medication use was also not systematically captured.
Some outcomes relied on medical billing and diagnostic codes
Electronic health records are valuable for large longitudinal studies, but ICD coding can misclassify or incompletely capture disease.
The authors report clinician review and high agreement in their adjudication process, but this remains a limitation.
External validation reduced the mortality effect
The independent cohort supported the basic finding, which is a major strength.
But the high-risk mortality hazard ratio fell from roughly 2.4 in the fully adjusted Cleveland analysis to about 1.6 in the external cohort.
That does not invalidate the model. It is exactly why independent validation matters.
The AI does not explain one simple causal mechanism
The system appears to integrate information from several physiological systems.
That may be precisely why it works.
It also makes it harder to translate a high-risk classification into a specific treatment.
Prospective trials have not shown clinical benefit
This is ultimately the most important limitation.
A risk predictor becomes medically valuable only if knowing the risk helps doctors make better decisions.
That still needs to be demonstrated.
Who funded the research?
Several authors reported research support from the Cleveland Clinic–IBM Discovery Accelerator Program, while other investigators reported National Heart, Lung, and Blood Institute funding.
Several authors were affiliated with IBM Research.
The published paper states that the authors declared no competing interests.
Industry involvement does not invalidate the results, particularly given publication in a peer-reviewed journal and external validation against an independent cohort.
It is nevertheless relevant context when evaluating technology that could eventually become a commercial clinical product.
So can a sleep study predict death?
The most accurate answer is:
A sleep study cannot currently predict that a specific person will die, much less determine when. But increasingly strong evidence shows that AI can extract patterns from overnight sleep physiology that are associated with future mortality and disease risk.
The August 2026 study adds an especially important finding.
The relevant information may already exist inside sleep studies doctors perform every day.
The problem is that conventional interpretation compresses an entire night of brain activity, cardiac signals, breathing, oxygenation, muscle activity and sleep architecture into a relatively small collection of summary numbers.
AI does not magically see the future.
It may simply be better at not throwing away as much of the present.
And if subsequent studies confirm that these hidden physiological patterns can guide treatment, the long-term implication could be much larger than an algorithm that “predicts death.”
A routine overnight sleep study could become a window into cardiovascular, neurological and overall physiological health—revealing risks that have been sitting in the data all along.
References and Further Reading
Primary research
-
A Foundation Model for Sleep-Based Risk Stratification and Clinical Outcomes — Nature Communications (2026) — The primary August 2026 study examining AI-derived physiological risk groups, mortality, cardiovascular outcomes and neurological outcomes in Cleveland Clinic sleep studies and an independent validation cohort.
-
A Multimodal Sleep Foundation Model for Disease Prediction — Nature Medicine (2026) — SleepFM study demonstrating prediction of mortality and numerous future diseases using a large multimodal polysomnography foundation model.
-
Age Estimation From Sleep Studies Using Deep Learning Predicts Life Expectancy — npj Digital Medicine (2022) — Earlier research showing that deep-learning-derived physiological age from sleep recordings was associated with mortality and estimated life expectancy.
Sleep testing and clinical context
-
Sleep Studies — National Heart, Lung, and Blood Institute — NIH overview of polysomnography and the physiological signals recorded during clinical sleep studies.
-
Clinical Practice Guideline for Diagnostic Testing for Adult Obstructive Sleep Apnea — American Academy of Sleep Medicine — Clinical guidance explaining sleep-apnea diagnostic testing and respiratory-event measures including AHI and RDI.
Consumer wearables
-
Receive Sleep Apnea Notifications on Apple Watch — Apple Support — Explains how Apple’s breathing-disturbance algorithm works and the limits of its sleep-apnea notification feature.
-
Oura & Medical Conditions — Oura Member Care — Oura’s current statement on the Ring’s medical limitations, including sleep apnea.
Editorial currency note: This article reflects research and consumer-device capabilities available as of August 20, 2026. AI sleep-risk models remain an active area of research, and clinical recommendations, device capabilities and regulatory approvals may change as prospective validation develops.



