OCT-Bench sets 10,076 questions to test whether multimodal models really understand retinal scans

Baochen Fu · hf · 2026-07-21

OCT-Bench introduces a 10,076-question benchmark for multimodal model understanding of OCT scans

The paper proposes OCT-Bench, a new benchmark for evaluating whether multimodal large language models can truly understand optical coherence tomography (OCT) images beyond coarse disease classification.

The benchmark is positioned as a more clinically grounded way to identify where multimodal models fail in medical image understanding.

Original post →

More from Multimodal

Multimodal channel →