The AI That Out-diagnosed Doctors in the Emergency Room
AI Versus Physicians: The Results Are In, And They’re Startling
The AI That Out-diagnosed Doctors in the Emergency Room
AI Versus Physicians: The Results Are In, And They’re Startling

A landmark study has put an advanced AI reasoning model head-to-head against human physicians, using real, messy, unprocessed hospital data. The AI won, and it wasn’t especially close. Could this be the moment medicine changes forever?
For more health and wellness information check out my YouTube Channel: https://www.youtube.com/@MyLongevityExperiment
Testing a “Thinking” Machine
The dream of a computer capable of diagnosing illness stretches back to at least 1959, when Ledley and Lusted published their landmark paper in Science, representing the beginning of the development of clinical decision support systems, interactive programs designed to assist physicians and other health professionals with decision making tasks. For decades, no software came close to matching the clinical depth of a trained physician.
That started changing with the rise of large language models, which ignited new hope and produced numerous studies with encouraging early results. The next major leap was the emergence of reasoning models, which maintain an internal chain of thought and can walk through their logic step by step before arriving at an answer. History of InformationInside Precision Medicine
That evolution has made the competition between humans and machines genuinely compelling. The first rigorous study pitting a reasoning large language model directly against human physicians has now been published in Science, authored by a team from Harvard Medical School and Beth Israel Deaconess Medical Center, with collaborators at Stanford. The model used was OpenAI’s first reasoning model, o1-preview, released in September 2024.
Despite the study being freshly published, the pace of AI development means this model is already considered outdated, and the newest generation should perform even better still. Harvard MagazineInside Precision Medicine
Outperforming Humans on Hard Cases
The research team tested o1-preview across six separate physician-style experiments, measuring it against hundreds of physicians and against the earlier GPT-4 model. In the first experiment, the researchers conducted five experiments to measure clinical reasoning across differential diagnosis generation, display of diagnostic reasoning, triage differential diagnosis, probabilistic reasoning, and management reasoning, all adjudicated by physician experts. arxiv
For the core diagnostic task, the team fed 143 NEJM clinicopathological conferences to the model and asked it to generate a ranked list of possible diagnoses. A CPC is a teaching format published regularly in the New England Journal of Medicine in which a faculty physician is asked to determine the diagnosis of an anonymous, actual patient based solely on medical history and preliminary test results.
These cases are typically complex and demanding, designed to stretch the reasoning abilities of experienced clinicians. PubMed Central
The AI reasoning model o1-preview outperformed all competing systems, including the correct diagnosis in its response almost 80% of the time. When answers judged to be very close were also counted, accuracy reached 97.9%. A legitimate concern with AI models working through published cases is the possibility of memorization, meaning the model may have encountered the case during training.
The authors addressed this directly by comparing performance on cases published before and after o1-preview’s training cutoff and found no meaningful difference, suggesting the model was genuinely reasoning rather than retrieving stored answers. Science News
GPT-4 performed noticeably worse across the board. More significantly, on a subset of 101 cases where human physician responses had previously been recorded, o1-preview surpassed those physicians on both its top single choice and across its broader ranked list.
What Should We Do Next?
Arriving at a diagnosis is only the beginning. A good clinician also needs to know what to do about it. To test this dimension, the team asked o1-preview which diagnostic test it would order next across 136 of the same cases. In 87.5% of cases, the model picked the correct test. In a further 11%, its choice was judged to be helpful, and only 1.5% of selections were deemed unhelpful. Inside Precision Medicine
The team then tested o1-preview on 20 cases drawn from NEJM Healer, a virtual patient educational tool, scoring responses across four domains of written clinical reasoning. On tasks involving what doctors refer to as “management reasoning,” from recommendations for antibiotic use to how to approach goals of care, including end-of-life conversations, o1-preview significantly outpaced previous AI models and also outperformed humans using conventional aids such as up-to-date Google search. Harvard Magazine
“Management reasoning is likely a more complex task than diagnostic reasoning,” explained Peter Brodeur, a clinical fellow at Beth Israel Deaconess Medical Center. “It requires many considerations of not only the objective features of a case, but also subjective factors: what context and situations you’re in, and therefore, it probably doesn’t come as a surprise that a reasoning model performs significantly better at such tasks than humans and ChatGPT-4.” Harvard Magazine
In one area, human physicians held their ground. The model was not meaningfully better than doctors at including the so-called “cannot-miss” diagnoses, the high-stakes possibilities that must remain on the table even when they seem remote.
To address memorization concerns even more rigorously, the authors used six diagnostic vignettes drawn from a 1994 study that have never been publicly released. o1-preview scored a median of 97% compared to 92% for GPT-4, 76% for physicians using GPT-4, and 74% for physicians using conventional resources, though the differences did not reach statistical significance given the small case count. Euronews
The Real-World Test
The element that distinguishes this study from earlier AI diagnostics research is its use of authentic, unprocessed clinical records.
Researchers based at Harvard Medical School and Beth Israel Deaconess Medical Center collected actual cases like those of patients who had visited the emergency department at Beth Israel in Boston and graded how well the AI model could provide an accurate diagnosis at three moments in time, from the triage stage in the ER, up to being admitted into the hospital. NPR
The team gathered 76 randomly selected emergency room cases with all identifiers and unstructured notes left completely intact. They constructed three diagnostic checkpoints: initial triage with minimal information available, the point of physician evaluation once history, physical exam, and initial labs were on hand, and finally admission to the ward or ICU when the most complete picture was available.
At each checkpoint, o1, GPT-4o, and two attending physicians independently produced differential diagnoses. Two separate attending physicians, kept blind to the source of each output, then scored every differential. Overall, AI outperformed two experienced physicians, and did so with only the electronic health records and the limited information that had been available to the physicians at the time. NPR
The blinding itself produced a striking finding. The raters identified whether a diagnosis came from AI or a human correctly only 3,15% of the time, choosing “Can’t tell” between 84,94% of the time. The model’s outputs were stylistically indistinguishable from those written by physicians.
The advantage was largest at the triage stage, where available data is thin and the stakes are highest. By the point of admission, when the full clinical picture was available, the gap narrowed and no longer reached statistical significance, suggesting that o1 extracts more diagnostic signal from sparse information than its human counterparts do.
“We didn’t pre-process the data at all,” said Adam Rodman, MD, MPH, a hospitalist and clinical researcher at BIDMC. “The model is literally just processing data as it exists in the health record.” “I thought it was going to be a fun experiment but that it wouldn’t work that well. That was not at all what happened.”
“We tested the AI model against virtually every benchmark, and it eclipsed both prior models and our physician baselines,” said co-senior author Arjun Manrai, assistant professor of biomedical informatics at Harvard’s Blavatnik Institute. “However, this does not mean AI will necessarily improve care, how and where it should be deployed remain understudied, and we desperately need rigorous prospective trials to evaluate the impact of AI on clinical practice.” Fortune
What This Doesn’t Mean
Caution is warranted before drawing sweeping conclusions. “A model might get the top diagnosis right but also suggest unnecessary testing that could expose a patient to harm,” Brodeur said. “Humans should be the ultimate baseline when it comes to evaluating performance and safety.” Euronews
Manrai emphasized that the team’s findings do not mean that “AI replaces doctors, despite what some companies selling AI-based healthcare are likely to say.” The study reflects model performance on historical and curated cases, not proof of safety or efficacy in live clinical settings.
The authors also noted that the study primarily focuses on the preview version of the o1 model, which has since been supplanted by newer models such as OpenAI’s o3, and that further studies are needed to understand how performance varies across models and how humans and LLMs may collaborate. Harvard MagazineEuronews
The broader implication is not replacement but transformation. The findings suggest that LLMs have now eclipsed most benchmarks of clinical reasoning, motivating the urgent need for human-computer interaction studies and prospective clinical trials to rigorously assess the potential of AI systems to improve clinical practice and patient outcomes, as the authors themselves put it.
The machine that physicians once laughed off as science fiction has, it seems, finally arrived at the clinic. Inside Precision Medicine
메타데이터
- post_id
- 6c2a6b4e06fd
- slug
- the-ai-that-out-diagnosed-doctors-in-the-emergency-room-6c2a6b4e06fd
- url
- https://medium.com/@mylongevityexperiment/the-ai-that-out-diagnosed-doctors-in-the-emergency-room-6c2a6b4e06fd
- canonical_url
- https://medium.com/@mylongevityexperiment/the-ai-that-out-diagnosed-doctors-in-the-emergency-room-6c2a6b4e06fd
- author_url
- https://medium.com/@mylongevityexperiment
- status
- ok
- fetched_at
- 2026-06-09 15:37:30