A study published April 30, 2026 in the journal Science, led by researchers at Harvard Medical School and Beth Israel Deaconess Medical Center with collaborators at Stanford, found that OpenAI’s o1 reasoning model outperformed practicing physicians across six separate diagnostic experiments — including on real emergency room cases pulled directly from patient charts, not just textbook scenarios built for testing.
Six experiments, one consistent result
The research team pitted o1 against hundreds of doctors spanning a wide range of experience levels: residents early in training, board-certified specialists, and family physicians who see undifferentiated symptoms all day. The tests ranged from classic symptom clusters long used in medical education to real-world data drawn from the charts of 76 patients who had actually visited a Boston-area emergency room. According to the authors, the model outperformed the human physicians in every one of the six experiments, without exception — a level of consistency that surprised even researchers who expected large language models to perform reasonably well on pattern-recognition-heavy diagnostic tasks.
The numbers behind the headline
In the experiment built around classic case studies, o1 generated a list of possible diagnoses that included the correct one 78% of the time, compared with roughly 30% for the physicians tested on the same cases. On the real emergency-room data, evaluated at the initial triage stage, o1 identified the exact or a very close diagnosis in 67.1% of cases, while the two human physicians in that comparison did so in 55.3% and 50% of cases, respectively. The gap held up whether the model was working from a tidy textbook vignette or messier real-world charting, which researchers say is the more clinically meaningful test since actual patients rarely present with cleanly organized symptom lists.
Why this is not simply “AI beats doctors”
The authors are careful to frame the result as narrower than a wholesale replacement of physicians. Diagnostic reasoning — generating a differential and narrowing it down from a description of symptoms and test results — is fundamentally a pattern-matching exercise, and it is exactly the kind of task large language models trained on enormous volumes of medical literature and case data are built to excel at. Other recent coverage of the same body of research has noted that while AI systems are now very good at diagnosing what is wrong with a patient, physicians remain distinctly better at the messier judgment calls that follow a diagnosis: weighing a patient’s individual circumstances, discussing treatment tradeoffs, and making decisions that depend on values and context a model cannot fully see.
What clinicians who reviewed the study are saying
Physicians who have examined the results broadly agree the findings track with what has been observed anecdotally in emergency departments already experimenting with AI-assisted triage: it is easier for a model to say “here is what this probably is” than to say “here is what we should do about it, given everything else going on with this patient.” That said, several clinicians quoted in coverage of the study have pushed back on framing this as evidence AI should operate unsupervised, noting the tested scenarios were retrospective and did not involve the model managing the ambiguity, incomplete information, and interpersonal dynamics of an actual live encounter with a patient in front of it.
The bigger question the study raises
The researchers themselves describe the results as evidence of an “urgent need” for controlled clinical trials that test these models prospectively, in real time, inside actual care settings rather than retrospectively against archived charts. That is a meaningfully higher bar than the six experiments already completed, and it is the bar regulators and hospital systems will likely require before diagnostic AI moves from a research curiosity to something deployed at the point of care. For now, the study adds to a growing body of evidence that large language models have crossed a real threshold in diagnostic pattern recognition — while leaving unresolved the much harder question of what role, if any, that capability should play in an exam room.
What comes next
Harvard and Beth Israel Deaconess researchers say follow-up work is already underway to test reasoning models prospectively against live clinical decision-making rather than retrospective chart review, and to examine whether the same performance gap holds across a broader range of specialties beyond emergency medicine. How fast hospitals move to adopt diagnostic AI as a formal second opinion — versus keeping it at arm’s length until prospective trials are complete — will likely shape how quickly this research translates into anything a patient actually encounters during a hospital visit.