Medmarks v1.0 lands NeurIPS track: medical LLM benchmark now covers 30 suites, 61 models
iScienceLuvr · x · 2026-09-26
Two updates for Medmarks, an open-source automated benchmark suite for evaluating LLM medical capabilities:
- Accepted to the NeurIPS Datasets and Benchmarks track (Sydney)
- v1.0 release with a technical report: benchmarks grew from 20 to 30, and models on the leaderboard from 46 to 61
The team bills it as the largest open-source automated medical LLM evaluation suite—useful reference for healthcare model selection and research.
Related event: Medical LLM Benchmark Medmarks v1.0 Accepted at NeurIPS(2 posts)→
More from Research
- Claude computes nine-loop scattering amplitude, breaking Lance Dixon's eight-loop record — kevinweil · 2026-09-26
- Stanford's Matryoshka Attribution tops mech interp benchmark, removes refusal by resetting 1% of weights — stanfordnlp · 2026-09-26
- Dual Covariance 3DGS SLAM: two covariances per Gaussian, robust tracking at 60 FPS — kwangmoo_yi · 2026-09-26
- Dual Covariance Gaussian Splatting SLAM decouples rendering and registration — kwangmoo_yi · 2026-09-26
- ImageJevBench: image decision benchmark ranks top models, full eval costs $0.02 — airesearch12 · 2026-09-26
- Biopharma Bench: agents complete only 8 of 71 real biopharma tasks, GPT-6 Astra leads — AllThingsApx · 2026-09-26