OCT-Bench sets 10,076 questions to test whether multimodal models really understand retinal scans
Baochen Fu · hf · 2026-07-21
OCT-Bench introduces a 10,076-question benchmark for multimodal model understanding of OCT scans
The paper proposes OCT-Bench, a new benchmark for evaluating whether multimodal large language models can truly understand optical coherence tomography (OCT) images beyond coarse disease classification.
- Scale: 10,076 multiple-choice questions built from 4,137 OCT images across 7 public datasets.
- Task design: a 20-task hierarchy spanning three levels — Perception, Cognition, and Reasoning.
- Coverage: imaging attributes, retinal anatomy, lesion characteristics, spatial relationships, disease assessment, treatment decisions, and prognosis management.
- Evaluation: 20 representative MLLMs were tested, including proprietary, open-source general-purpose, and medical-domain models.
- Result: current models are still far from reliable OCT understanding; neither medical adaptation nor larger scale consistently improves performance across capability levels.
The benchmark is positioned as a more clinically grounded way to identify where multimodal models fail in medical image understanding.
More from Multimodal
- Fable made the music, the instrument, and the video for The Loom — Sauers_ · 2026-07-21
- Open-source TTS list sorts models by license before quality for commercial shipping — mahimairaja · 2026-07-21
- AI pastiche turns Chris Cornell’s “American Nightmare” into a meme poster — 77sevens · 2026-07-21
- AI gaming demos now generate real-time sound to match world-model video — mark_k · 2026-07-21
- Krea2 users find a 4-step Raw plus 4-step Turbo workflow that preserves quality — PropagandaOfTheDude · 2026-07-21
- Hugging Face screenshot shows Anima diffusion model versions including Aesthetic v1.1 — RafiHDW · 2026-07-21