MMBU Challenge details: three tracks, API evals for 27B+ models, adaptation scored by lift
LiorOnAI · x · 2026-10-01
The MMBU (Massive Multimodal Biomedical Understanding) Challenge published its rules, testing whether multimodal models actually perceive biomedical images rather than lean on priors.
- Core thesis: visual perception is the primary bottleneck for reliable biomedical MLLM reasoning
- One checkpoint per track; models over 27B active parameters must be evaluated via API
- Track 1 Open Frontier: best overall, no size cap, all training stages from one org, scored 0.8×task + 0.2×biomedical context
- Track 2 Domain Adaptation: adapt a disclosed base model toward medicine, ranked by accuracy lift over the base, not starting strength
- Backed by Anthropic, Stanford AI Lab, biohub; runs Oct 1–Dec 31
Related event: Stanford and Partners Launch $100K MMBU Biomedical Multimodal Challenge(3 posts)→
More from Models
- 19 fixtures, 6 models, 3 runs: newest AI models didn't beat the old one at finding bugs — ChanceKelch · 2026-10-01
- Kilpatrick: Gemini 4 Argon is 'just the start' of Google's model progress — OfficialLoganK · 2026-10-01
- Google launches Gemini 4 Argon with 1M-token output, $2/$10 intro pricing — _philschmid · 2026-10-01
- Google previews Gemini 4 Argon benchmarks, rolls out to cyber defenders first — ammaar · 2026-10-01
- Gemini 4 Argon ships with industry-leading 1M token output limit, frontier long-horizon reasoning — GoogleAI · 2026-10-01
- Gemini 4 Argon said to hit SoTA, with claims of saving 300TB memory in data centers — thesaraharminta · 2026-10-01