ECCV MMBU Benchmark Shows VLMs Answer Biomedical Questions Without Knowing What They See
davidjhwu · x · 2026-09-24
Frontier vision-language models often give convincing answers to biomedical questions while failing to identify the modality, body part, or stain in the image, according to the MMBU benchmark presented at ECCV. The MMBU Challenge is now inviting teams to tackle this perception bottleneck: three tracks, $100K+ in compute and prizes, running Oct 1–Dec 31, with registration extended through September 30. Supported by gxlai, Anthropic, Stanford AI Lab and others.
More from Multimodal
- Not the tech, the terms: a Hollywood veteran's century-spanning case for how AI reshapes creative work — ccerrato147 · 2026-09-24
- Meta's Muse video model tipped to be wildly popular, but users balk at handing data to Zuck — kyliebytes · 2026-09-24
- MiniMax H3's Sailor Moon Clip Gets Weirdly Random Zoom-Ins — Certain_Potato_4509 · 2026-09-24
- MiniMax H3 video VAE gets ~2.2x faster encoding after Kijai's optimization — Lexius2129 · 2026-09-24
- WVLNGTH Brighton puts 9 AI video creators' work on cinema screens — LudovicCreator · 2026-09-24
- Pixel Art Test Improved by Having GPT-6-Sol Use PixiJS Instead of Pure JS — burny_tech · 2026-09-24