Nature Medicine Study: LLMs Excel at Triage Discrimination But Lack Calibration
erichorvitz · x · 2026-08-04
A recent Nature Medicine study highlighted by Eric Horvitz reveals that high-reasoning LLMs demonstrate remarkable discrimination in emergency triage, achieving AUROCs between 0.95 and 0.99 for ranking patients by urgency.
However, the models often suffer from poor calibration. Addressing previous findings where ChatGPT under-triaged emergencies, the analyses suggest a nuanced explanation: the models generally recognize relative danger but struggle with absolute probability thresholds.
More from Research
- Hidden LLM prompts can be reverse-engineered from outputs alone — mhmazur · 2026-08-05
- Biomni Lab Partners with Chugai to Deploy AI Drug Discovery Platform — KexinHuang5 · 2026-08-05
- Proposing 'Latent Reading': Exploring Literature via Single-Text AI Models — begusgasper · 2026-08-04
- Practical Alignment Agenda: Eradicating Reward Hacking and Model Deception — MariusHobbhahn · 2026-08-04
- Why LLMs Corrupt Your Notes and How to Fix It Structurally — Cryvixx · 2026-08-04
- Frontier VLMs Fail at Jigsaw Puzzles: JigShape Benchmark Tests Geometric Reasoning — _vztu · 2026-08-04