ICDAR 2026 Competition: Multimodal AI Struggles with Scientific Figures
SciKnowOrg · hf · 2026-08-04
ICDAR 2026 hosted a competition on information extraction from ALD/E scientific figures, releasing the Sci-ImageMiner benchmark.
- Challenge & Scale: Requires integrating visual perception with domain-specific reasoning. The competition attracted 68 participants and 1,263 submissions.
- Model Performance: SOTA multimodal models perform well on classification/summarization but struggle with data extraction and scientific reasoning (VQA).
The findings highlight key limitations of current AI in domain-specific figure comprehension.
More from Research
- ARR August Cycle Receives 4,134 Submissions, Up 14% Year-over-Year — mdlhx · 2026-08-04
- Nature Publishes AI Biologist XunZi for Disease-Modifying Target Discovery — EricTopol · 2026-08-04
- Ai2 Launches OLMoEarth: Open-Source Planetary-Scale Geospatial Inference Platform — anselm · 2026-08-04
- Nature Study: AI Dermatology Diagnosis Amplifies Public Automation Bias — EricTopol · 2026-08-04
- Radical Co-founder on Training the Largest Genome Model to Write DNA — exnx · 2026-08-04
- 3 Lines of Code Fixed 123 Failed PPO Experiments by Changing Reward Shaping — mikeysce · 2026-08-04