LLM judges verify presence, not absence: study reveals omission blindness in AI clinical notes
sbulaev · hn · 2026-09-02
An arXiv paper documents "omission blindness" in LLM-as-judge pipelines for AI-generated clinical notes: LLM reviewers are good at verifying that what's present is correct, but systematically fail to catch critical omissions—missing symptoms, unrecorded medications, omitted abnormal findings. The authors argue this is the most dangerous failure mode in medical note review, since a wrong statement is more likely to be caught than a missing one, and present evaluation findings analyzing when and why judge models miss absences.
More from Research
- Burkov skew AI hype: 'deterministic LLMs' and 'first agents' are old tricks rebranded — burkov · 2026-09-23
- Continuous diffusion beats discrete on random k-SAT, proposed as standard benchmark — ArashVahdat · 2026-09-23
- Grady Booch: Contemporary AI Still Lacks Abductive Reasoning, Just 'Next-Token Prediction' — Grady_Booch · 2026-09-23
- AI solves Navier-Stokes-related problem as machines upend mathematics, New Scientist reports — burny_tech · 2026-09-23
- Mathematician says OpenAI likely proved a significant partial case of the Hodge conjecture — burny_tech · 2026-09-23
- Code benchmarks are mostly slop: dev calls for narrow evals per domain, not one score — almmaasoglu · 2026-09-23