Medmarks Upgraded to v1.0
iScienceLuvr · x · 2026-07-11
SophontAI has released Medmarks v1.0 alongside a technical report, marking an update to this open-source automated benchmark suite for evaluating LLM medical capabilities.
Updates include:
- Added 10 new benchmarks, bringing the total from 20 to 30
- Added 15 new models, expanding leaderboard coverage from 46 to 61
- Presented a poster at the ICML Large Language Models for Life Sciences workshop
More from Research
- WeirdChat catalogs strange model behaviors from more than 100 million sampled responses — JacobSteinhardt · 2026-07-22
- New agentic benchmark shows AI managers escalate to coercion and fake success — Jasmine Brazilek · 2026-07-22
- Ai2’s Asta adds one-click handoff and self-checking deep paper search — allen_ai · 2026-07-22
- NVIDIA says physical AI starts in simulation with OpenUSD and synthetic data — MonaJalal_ · 2026-07-22
- DepthART pushes monocular depth to tiny models at 1000 FPS on RTX A6000 — kwangmoo_yi · 2026-07-22
- Meta says SAM 3 and DINOv3 cut 3D volume labeling from a month to 15 minutes — AIatMeta · 2026-07-22