NeurIPS GenAI4Health oral: retrieval-based medical fact-checking fails in ways bigger models can't fix
mdredze · x · 2026-10-09
A paper from Mark Dredze's group at Johns Hopkins, "Where Does Retrieval-Based Open-Ended Evaluation Fail?" (led by Heyuan Huang), was accepted as an Oral at NeurIPS GenAI4Health (top 4% of 250 submissions). The team also won Best Abstract (People's Choice) at DEX26 for a benchmark showing top LLMs catch many physician diagnostic errors but big gaps remain.
The paper tackles the dominant retrieve-then-verify paradigm for medical hallucination detection: aggregate metrics like F1 hide where and why systems fail. The authors build two taxonomies — retrieval-stage errors along five quality dimensions, and verifier-reasoning errors across six consecutive steps — using an LLM-as-Judge pipeline to label evidence quality and classify verifier errors at scale, then stress-test across 4 retrieval methods and 6 frontier verifiers.
Key finding: scaling model size, adding reasoning effort, expanding to authoritative web sources, and medical fine-tuning do not fix these systematic failures.
More from Research
- Inherit-MAS cuts multi-agent token use by up to 34.6% with evolution-inspired inheritance — Songtao Wei · 2026-10-09
- Zero human labels: auto-generated soccer tracking dataset hits 62.5 HOTA, beating fine-tuned FairMOT — RexDouglass · 2026-10-09
- Why AI doesn't actually read words: from BPE subwords to byte-level models like BLT and H-Net — jbhuang0604 · 2026-10-09
- NVIDIA details HSTU recommender inference stack with up to 5.93x lower latency — PyTorch · 2026-10-09
- Preprint: LLMs store numbers as curves and helices, but compute comparisons differently — tweetsatpreet · 2026-10-09
- Best model was cheapest: open-weights model ran 669 clinical decisions for 1.7 cents — antoine_chaffin · 2026-10-09