Medmarks update: Gemma 4 31B leads mid-size open-source models on medical benchmark

iScienceLuvr · x · 2026-09-21

Sophont updated Medmarks, its open-source medical LLM benchmark suite and leaderboard, adding results for Gemma, Qwen, Muse, and Nemotron in the 20-40B range. Gemma 4 31B leads its size class; Qwen 3.8 27B can't match Sonnet on health tasks. The suite has evaluated 70 models across 89 configs, with public datasets, prompts, and grading code — scores are relative win rates, reproducible on local hardware.

Related event: Medmarks Benchmark Update: Gemma 4 31B Tops Mid-Size Open Models(2 posts)→

Original post →

More from Models

Models channel →