Stale benchmarks plague medical LLM evals; new living benchmark for EHR retrieval

davidjhwu · x · 2026-10-06

Researchers highlight a big problem in medical benchmarking: benchmarks go stale. A new preprint introduces "A Living Benchmark" for evaluating whether evolving LLMs reliably retrieve information clinicians need from electronic health records. A notable finding: omission is a common failure mode, echoed by the NOHARM preprint team.

Original post →

More from Research

Research channel →