Uncategorized

Mount Sinai Study Finds ChatGPT Health Missed More Than Half of Cases Needing Emergency Care

A Nature Medicine study found OpenAI's ChatGPT Health under-triaged more than half of emergency-care cases and showed inverted suicide-risk alerts, testing the tool across 960 clinical interactions.

Mount Sinai Study Finds ChatGPT Health Missed More Than Half of Cases Needing Emergency Care

A study fast-tracked into the February 23, 2026 online issue of Nature Medicine found that ChatGPT Health, OpenAI’s consumer health-guidance feature launched in January 2026, under-triaged more than half of test cases that physicians determined required emergency care. The research, led by Ashwin Ramaswamy, MD, and senior author Girish N. Nadkarni, MD, MPH, of the Icahn School of Medicine at Mount Sinai, is described as the first independent safety evaluation of the tool since its launch.

How the study was built

Researchers constructed 60 structured clinical scenarios spanning 21 medical specialties, then layered in 16 contextual variations per scenario — changing details like the patient’s race, gender, social circumstances, and access barriers to care — producing 960 total interactions with ChatGPT Health. Three independent physicians, working from 56 medical society guidelines, determined the clinically correct urgency level for each scenario before comparing it against what the AI tool actually recommended.

What went wrong, specifically

The tool performed reasonably well on textbook emergencies with unambiguous symptoms, according to the study, but struggled badly with nuanced or atypical presentations — the kind of cases where clinical judgment matters most in deciding whether someone needs to go to an emergency room immediately or can wait for a regular appointment. Most alarmingly, the researchers found that ChatGPT Health’s safety alerts for suicide risk were, in their words, effectively inverted relative to clinical risk: alerts triggered more reliably for lower-risk scenarios than for cases in which a person described a specific plan to hurt themselves. “LLMs have become patients’ first stop for medical advice — but in 2026 they are least safe at the clinical extremes, where judgment separates missed emergencies from needless alarm,” one of the study’s lead investigators said.

Why this matters at ChatGPT Health’s scale

ChatGPT Health had amassed roughly 40 million daily users seeking medical guidance within weeks of its January 2026 launch, according to figures reported alongside the study, ranging from routine symptom questions to urgent triage decisions about whether to seek emergency care. That scale is the crux of the concern: a triage error rate that would be worrying in a small pilot becomes a public health issue when tens of millions of people are relying on the same system daily, often as a substitute for calling a nurse line or a doctor’s office.

OpenAI’s position and the broader debate

OpenAI has positioned ChatGPT Health as a guidance tool rather than a replacement for professional medical care, and the product includes disclaimers directing users to seek emergency services for serious symptoms. Supporters of consumer health AI argue that even an imperfect triage tool may outperform the alternative for the many people who have no easy access to a nurse hotline or primary care physician, and that the alternative to an AI tool with flaws is often no guidance at all. Critics, including patient-safety advocates cited in coverage of the study, counter that a triage tool used by tens of millions of people needs a far higher reliability bar than a typical consumer product, particularly around suicide risk, where the stakes of a missed signal are irreversible.

Where this fits the AI-triage debate more broadly

The Mount Sinai findings arrive alongside a parallel, more favorable body of research on AI triage tools deployed inside hospitals under clinical supervision — such as Yale New Haven Health’s emergency department triage system, which has shown improved identification of critical patients when overseen by trained staff. The contrast underscores a distinction researchers are increasingly drawing: AI triage tools embedded in a clinical workflow with human oversight appear to behave very differently than consumer-facing chatbots operating with no clinician in the loop.

What happens next

The Nature Medicine study is likely to intensify calls from medical societies and patient-safety groups for regulatory scrutiny of consumer AI health tools that give triage-adjacent advice, an area that currently falls outside FDA’s traditional medical device framework because these tools are marketed as general information services rather than diagnostic devices. Mount Sinai researchers say further work is needed to test whether targeted fixes to suicide-risk detection and edge-case triage can close the gap before consumer AI health tools reach even larger audiences. Nadkarni and his co-authors say they plan to share their scenario-testing methodology with other research groups so that future consumer health AI products can be benchmarked before, not after, they reach tens of millions of users.