Study Maps Why Retrieval-Based Medical Factuality Evaluation Fails — Bigger Models Won't Fix It
jhu-clsp · hf · 2026-10-06
A new study builds two taxonomies for failures in retrieval-based factuality verification of open-ended medical answers: retrieval-stage errors across five quality dimensions and verifier-reasoning errors across six consecutive steps, labeled at scale with an LLM-as-Judge pipeline.
- Stress-tested across 4 retrieval methods and 6 frontier verifier models on MedExpert and three closed-ended datasets
- Scaling model size, adding reasoning effort, expanding authoritative web sources, and medical fine-tuning all fail to resolve the failure modes
- Conclusion: these are fundamental limits of the retrieve-then-verify paradigm in open-ended medical settings, not artifacts of outdated systems
- Code and data released for reproducibility
More from Research
- Subsampling and extrapolation keep the Mandelbrot area estimate unbiased near the boundary — geoffreyirving · 2026-10-06
- Mandelbrot area estimate builds on Böttcher-series upper bound, open-sourced with cross-checked precision — geoffreyirving · 2026-10-06
- New estimate sits 6.5e-9 below Hsing Lo's 2025 value; reproduction suggests the gap is a fluctuation — geoffreyirving · 2026-10-06
- Claude-assisted CUDA compute pins Mandelbrot set area to 1.506591883653, 60x tighter than 2012 record — geoffreyirving · 2026-10-06
- How the new Mandelbrot area estimate works: quadtree pruning plus random boundary sampling on H200s — geoffreyirving · 2026-10-06
- DeepMind's AlphaProtein Novo designs new-to-nature enzymes, but weights stay closed — anshulkundaje · 2026-10-06