Even as OpenAI pushes its healthcare products deeper into major U.S. hospital systems, a new independent study out of Mount Sinai is raising sharp questions about whether the underlying models are ready for the stakes involved. The study, led by Ramaswamy and colleagues and published in Nature Medicine in February 2026, evaluated ChatGPT Health’s performance on emergency triage scenarios and found troubling failure rates on exactly the cases where getting it wrong is most dangerous.
The Numbers Behind the Warning
The Mount Sinai team found that ChatGPT Health under-triaged 52 percent of gold-standard emergency presentations, meaning cases that clinical experts agreed required urgent or emergency care were instead flagged by the AI as less severe. In the opposite direction, the model over-triaged 35 percent of non-urgent presentations, incorrectly escalating cases that did not need emergency attention. Researchers also documented a significant anchoring bias: when a patient’s message minimized their own symptoms or downplayed distress, the model tended to follow that framing rather than probe further, even when the underlying symptoms described a genuine emergency.
A Gap in the Safety Net
Perhaps most concerning to the study authors was inconsistent activation of suicide-crisis safeguards, the automated prompts and crisis-line referrals that AI health tools are supposed to trigger when a conversation suggests a mental health emergency. The inconsistency suggests that safety guardrails built for one style of user message may not generalize reliably across the varied, often indirect ways real patients describe distress.
OpenAI’s Parallel Expansion
The findings land at an awkward moment for OpenAI’s healthcare strategy. The company launched OpenAI for Healthcare in January 2026, a suite of HIPAA-compliant products built on its GPT-5.2 models, and says the tools are already deployed across health systems including AdventHealth, Cedars-Sinai, HCA Healthcare, Memorial Sloan Kettering Cancer Center, Stanford Medicine Children’s Health and UCSF. OpenAI has also built its own evaluation tools, HealthBench and the newer HealthBench Professional, developed with 262 physicians across 60 countries and graded against more than 48,000 clinician-written rubric criteria, to argue that its models are improving on real clinical benchmarks.
Competing Claims, Different Yardsticks
OpenAI’s internal benchmarks and Mount Sinai’s independent study are measuring different things and reaching different conclusions, which is itself part of the problem industry watchers are flagging. HealthBench evaluates model responses against physician-written rubrics for accuracy, completeness and communication in synthetic multiturn conversations. The Mount Sinai study instead tested real triage decision-making against a gold standard of actual emergency severity, a harder and arguably more consequential test since a triage error can directly delay or accelerate care. The gap between a model performing well on a written rubric and performing safely on live triage decisions is exactly what critics of rapid AI healthcare deployment have been warning about.
The Bigger Context
The tension fits into a broader pattern this year: healthcare AI adoption is accelerating rapidly, with UnitedHealth reporting billions in projected savings from AI deployment and Anthropic reportedly positioning its own healthcare and biology work as central to its IPO pitch, even as independent researchers keep surfacing safety gaps that internal corporate benchmarks did not catch. A separate industry poll this year found governance and clinician trust remain the top unresolved barriers to scaling hospital AI, a finding that lines up uncomfortably well with what Mount Sinai’s triage data suggests.
What Happens Next
Expect regulators and hospital risk-management teams to lean harder on independent, adversarial testing like Mount Sinai’s rather than relying solely on vendor-published benchmarks before expanding patient-facing AI triage tools. OpenAI has not disputed the Nature Medicine findings publicly but continues to expand ChatGPT for Clinicians, a free tool aimed at physicians rather than patients directly, which sidesteps some of the direct-to-consumer triage risk the study highlights. Whether that distinction, AI assisting a clinician versus AI advising a patient directly, becomes the dividing line for future regulation is likely to be the central healthcare AI policy fight of the next year.