Nature Medicine Study: LLMs Excel at Triage Discrimination But Lack Calibration

erichorvitz · x · 2026-08-04

A recent Nature Medicine study highlighted by Eric Horvitz reveals that high-reasoning LLMs demonstrate remarkable discrimination in emergency triage, achieving AUROCs between 0.95 and 0.99 for ranking patients by urgency.

However, the models often suffer from poor calibration. Addressing previous findings where ChatGPT under-triaged emergencies, the analyses suggest a nuanced explanation: the models generally recognize relative danger but struggle with absolute probability thresholds.

Related event: Study Reveals LLMs Excel in Medical Triage but Suffer from Utility Misalignment(2 posts)→

Original post →

More from Research

Research channel →