Medmarks adds results for mid-size open LLMs, Gemma 4 31B leads the class
iScienceLuvr · x · 2026-09-21
Sophont released additional Medmarks v1.0 results for recent mid-size open-source LLMs including Gemma, Qwen, Muse, and Nemotron, with Gemma 4 31B leading this size class. Medmarks is an open benchmark suite and leaderboard for LLM medical capabilities. Maintainer Benjamin Warner also teased that the next release will evaluate more models, including ones codenamed Astra and Fable.
Related event: Medmarks Benchmark Update: Gemma 4 31B Tops Mid-Size Open Models(2 posts)→
More from Models
- CursorBench 4.0: Grok 4.7 xhigh hits 46.3% at $6.01/task, big jump over 4.6 at same cost — haider1 · 2026-09-22
- Grok 4.7 cuts hallucination rate to 29% from 34% on AA-Omniscience — ArtificialAnlys · 2026-09-22
- Same-prompt test: Grok 4.7 takes 32 min at $8.14 while free SWE 2 finishes in 15 min — iamfakhrealam · 2026-09-22
- Grok 4.7 measured at ~188 tokens/second, ~7.1 minutes per Intelligence Index task — ArtificialAnlys · 2026-09-22
- Grok 4.7's 81k output tokens per task more than double Grok 4.6's 36k — ArtificialAnlys · 2026-09-22
- Grok 4.7 jumps on coding index but burns 81k tokens per task, 2x its predecessor — ArtificialAnlys · 2026-09-22