AI diagnosed ER patients more accurately than two physicians in 67% of triage cases. It also recommended unnecessary tests that could do more harm than good.
“We tested the AI model against virtually every benchmark, and it eclipsed both prior models and our physician baselines.” — Arjun Manrai, Harvard Medical School
“AI is good at diagnosing, but it also tends to suggest unnecessary testing that could actually do more harm than good.” — Peter Brodeur, Beth Israel / Harvard
Published in Science, this Harvard Medical School study tested OpenAI’s o1 model against two attending physicians on 76 real emergency cases from Beth Israel Deaconess Medical Center in Boston. The AI received identical information to the doctors with no pre-processing: raw electronic health records as they appeared at the time of each diagnosis. Blind reviewers (two additional physicians who did not know which answers came from AI or humans) found o1 gave the exact or near-correct diagnosis in 67% of triage cases, compared to 55% and 50% for the two physicians. With more complete data (labs, imaging), accuracy rose to 82% for AI versus 70-79% for doctors. On management reasoning (antibiotic decisions, end-of-life care), AI scored 89% against a 46-physician baseline of 34%. However, the study found AI tends to recommend unnecessary tests, and the researchers explicitly stated the results do not mean AI is ready for live clinical decisions, calling for prospective trials before any deployment near actual patients.
accuracy
accuracy
reasoning score
(46 doctors)
Superior diagnostic accuracy does not mean safe for deployment. That determination requires independent evaluation of the complete behavioral profile, not just the accuracy metric. The same AI that outperformed physicians on diagnosis also recommended unnecessary tests that could harm patients. The researchers themselves said the model is not ready for clinical decisions. This is exactly the gap AVAAS fills: independent certification evaluates not just whether the AI gets the right answer, but whether it does harm along the way. A model that diagnoses correctly but orders unnecessary invasive procedures has a harm-of-inaction and irreversibility profile that accuracy benchmarks alone cannot capture. AVAAS measures the full behavioral surface, not just the headline metric.
Buckley, T., Brodeur, P., Manrai, A. et al. (May 2026). Published in Science. Covered by: Harvard Magazine, TechCrunch, NPR, Fortune
This entry is one of 37 documented cases in the AVAAS evidence ledger, a public record of AI and automated-system failures with a verified source on every entry.
Every case here reached a person.
AVAAS certifies how AI systems behave at the decision point, with documented third-party evidence of conformity to a published standard.
Certify Your AI →