OCT-Bench sets 10,076 questions to test whether multimodal models really understand retinal scans
Baochen Fu · hf · 2026-07-21
OCT-Bench introduces a 10,076-question benchmark for multimodal model understanding of OCT scans
The paper proposes OCT-Bench, a new benchmark for evaluating whether multimodal large language models can truly understand optical coherence tomography (OCT) images beyond coarse disease classification.
- Scale: 10,076 multiple-choice questions built from 4,137 OCT images across 7 public datasets.
- Task design: a 20-task hierarchy spanning three levels — Perception, Cognition, and Reasoning.
- Coverage: imaging attributes, retinal anatomy, lesion characteristics, spatial relationships, disease assessment, treatment decisions, and prognosis management.
- Evaluation: 20 representative MLLMs were tested, including proprietary, open-source general-purpose, and medical-domain models.
- Result: current models are still far from reliable OCT understanding; neither medical adaptation nor larger scale consistently improves performance across capability levels.
The benchmark is positioned as a more clinically grounded way to identify where multimodal models fail in medical image understanding.
More from Multimodal
- FLUX.2 Klein Drifts Hard on Character Expressions While Free Gemini Holds Likeness — wacomlover · 2026-09-11
- Tencent Hunyuan releases AuK code and weights on GitHub with ComfyUI and fine-tuning support — aigclink · 2026-09-11
- Creator turns Bahamut vs Tiamat rivalry into an AI cinematic battle with Midjourney, GPT Image 2 and Seedance — azed_ai · 2026-09-11
- invideo launches AI agent-powered editor to automate repetitive editing tasks — azed_ai · 2026-09-11
- fable 5.1 recreates The Starry Night with 256,157 JavaScript brush strokes — cedric_chee · 2026-09-11
- GPT-6 Astra + Hyper3D Rodin MCP Generates 3D Assets in One Agent Flow — ahuja_priyank · 2026-09-11