Blind-Spots-Bench for Multimodal Models Released

epfl-ch · hf · 2026-07-15

Blind-Spots-Bench aims to evaluate the "blind spots" of multimodal models: tasks that humans find trivial but where current models frequently fail. ## What They Did - Collected and cleaned raw questions from AI course students, curating them into **235** samples. - Supplemented the dataset with structured reference answers and designed a task taxonomy tailored to it. - Built an automated scoring pipeline covering language models, vision-language models, and image generation models, including both open-weight and closed-source models. ## Key Findings - Closed-source frontier models significantly outperformed open-source models overall by a margin of roughly **10%**, even when their scores on traditional benchmarks were comparable. - No single model dominates across all task types. - Certain tasks remained difficult for all tested models. The authors conclude that such "blind spot" benchmarks are better at diagnosing the specific weaknesses of current models, rather than just looking at standard benchmark scores.

Original post →

More from Multimodal

Multimodal channel →