NeurIPS GenAI4Health oral: retrieval-based medical fact-checking fails in ways bigger models can't fix

mdredze · x · 2026-10-09

A paper from Mark Dredze's group at Johns Hopkins, "Where Does Retrieval-Based Open-Ended Evaluation Fail?" (led by Heyuan Huang), was accepted as an Oral at NeurIPS GenAI4Health (top 4% of 250 submissions). The team also won Best Abstract (People's Choice) at DEX26 for a benchmark showing top LLMs catch many physician diagnostic errors but big gaps remain.

The paper tackles the dominant retrieve-then-verify paradigm for medical hallucination detection: aggregate metrics like F1 hide where and why systems fail. The authors build two taxonomies — retrieval-stage errors along five quality dimensions, and verifier-reasoning errors across six consecutive steps — using an LLM-as-Judge pipeline to label evidence quality and classify verifier errors at scale, then stress-test across 4 retrieval methods and 6 frontier verifiers.

Key finding: scaling model size, adding reasoning effort, expanding to authoritative web sources, and medical fine-tuning do not fix these systematic failures.

Related event: NeurIPS Paper Shows Systemic Failure in Medical Retrieval-Based Fact-Checking(3 posts)→

Original post →

More from Research

Research channel →