MedExpert dataset: 540 expert-annotated pairs to evaluate medical chatbot reliability
mdredze · x · 2026-09-18
Mark Dredze's group at Johns Hopkins congratulates postdoc Sonal Joshi on completing her postdoc. Their work focused on systems that identify errors and omissions in medical chatbot responses, with one paper published and more on the way.
The published work is MedExpert, a dataset for evaluating medical chatbots, presented at the Fifth Machine Learning for Health Symposium (PMLR 297).
- 540 question-response pairs across two specialties: young adult mental health and prenatal care
- Each response annotated by clinical subject-matter experts for factual accuracy, completeness, and more
- Targets the reliability concerns of patient-facing LLM medical chatbots in clinical settings
- Provides a framework for evaluating automatic error-detection systems in these domains
More from Research
- MoDA: RL alignment method fights LLM mode collapse while preserving output quality — stanfordnlp · 2026-09-19
- Stanford study finds the brain is actually two separate organs working together — Dr_Singularity · 2026-09-19
- Apple researchers propose probe guidance, cutting guidance cost for diffusion LMs with no extra forward pass — itsbautistam · 2026-09-19
- New Science paper shows disorder and heterogeneity can stabilize complex networks — wgilpin0 · 2026-09-19
- Token Superposition Training Cuts Pretraining Compute 2.5x on 10B MoE — gordic_aleksa · 2026-09-19
- First-of-its-kind AI x Med Chem Hackathon in Boston Ends; Compounds Head to Synthesis — generativist · 2026-09-19