Stale benchmarks plague medical LLM evals; new living benchmark for EHR retrieval
davidjhwu · x · 2026-10-06
Researchers highlight a big problem in medical benchmarking: benchmarks go stale. A new preprint introduces "A Living Benchmark" for evaluating whether evolving LLMs reliably retrieve information clinicians need from electronic health records. A notable finding: omission is a common failure mode, echoed by the NOHARM preprint team.
More from Research
- Huawei Noah's Tail-Influence Sampling Cuts CVaR Policy Evaluation MSE by Up to 76% — huawei-noah · 2026-10-06
- Google's KeyRec Achieves Best Long-Video VLM Results With Just 10% of Visual Token Budget — google · 2026-10-06
- 4DCodeBench Shows Frontier Models Reconstruct Static Scenes but Fail at Dynamics — 4DCodeBench · 2026-10-06
- OmniTaskonomy: Year-long study shows generation training can improve understanding tasks — WeijiaShi2 · 2026-10-06
- Newton proved the product rule without limits, using a discrete symmetric-difference trick — ctjlewis · 2026-10-06
- DeepMind's AI designs enzymes from scratch: 99x drug building block yield, plastic-eating at 90°C — 141_1337 · 2026-10-06