Medmarks update: Gemma 4 31B leads mid-size open-source models on medical benchmark
iScienceLuvr · x · 2026-09-21
Sophont updated Medmarks, its open-source medical LLM benchmark suite and leaderboard, adding results for Gemma, Qwen, Muse, and Nemotron in the 20-40B range. Gemma 4 31B leads its size class; Qwen 3.8 27B can't match Sonnet on health tasks. The suite has evaluated 70 models across 89 configs, with public datasets, prompts, and grading code — scores are relative win rates, reproducible on local hardware.
Related event: Medmarks Benchmark Update: Gemma 4 31B Tops Mid-Size Open Models(2 posts)→
More from Models
- Azure OpenAI content filter blocks 'S&M' — the standard finance shorthand for Sales & Marketing — peterjliu · 2026-09-22
- IFM's K2-Horizon-36B-A4B Matches 20x-Larger Models on AA Index Using New MoVA Architecture — victormustar · 2026-09-22
- Grok 4.7 fails again: $1.59 run produces laughable output — teortaxesTex · 2026-09-22
- LLM scam detection benchmarked: fitted TF-IDF baseline beats Jev, DeepSeek and local Qwen — justinbiebar · 2026-09-22
- Goodfire Finds DNA Model Evo 2 Encodes the Tree of Life as a Curved Activation Manifold — burny_tech · 2026-09-22
- Internal eval puts Grok 4.7 at #3 across 22 knowledge-work tasks for under $5 — realsohamparekh · 2026-09-22