ECCV MMBU Benchmark Shows VLMs Answer Biomedical Questions Without Knowing What They See

davidjhwu · x · 2026-09-24

Frontier vision-language models often give convincing answers to biomedical questions while failing to identify the modality, body part, or stain in the image, according to the MMBU benchmark presented at ECCV. The MMBU Challenge is now inviting teams to tackle this perception bottleneck: three tracks, $100K+ in compute and prizes, running Oct 1–Dec 31, with registration extended through September 30. Supported by gxlai, Anthropic, Stanford AI Lab and others.

Original post →

More from Multimodal

Multimodal channel →