Blind-Spots-Bench for Multimodal Models Released
epfl-ch · hf · 2026-07-15
Blind-Spots-Bench aims to evaluate the "blind spots" of multimodal models: tasks that humans find trivial but where current models frequently fail. ## What They Did - Collected and cleaned raw questions from AI course students, curating them into **235** samples. - Supplemented the dataset with structured reference answers and designed a task taxonomy tailored to it. - Built an automated scoring pipeline covering language models, vision-language models, and image generation models, including both open-weight and closed-source models. ## Key Findings - Closed-source frontier models significantly outperformed open-source models overall by a margin of roughly **10%**, even when their scores on traditional benchmarks were comparable. - No single model dominates across all task types. - Certain tasks remained difficult for all tested models. The authors conclude that such "blind spot" benchmarks are better at diagnosing the specific weaknesses of current models, rather than just looking at standard benchmark scores.
More from Multimodal
- Creator makes a dark-fantasy short film teaser with Google Flow visuals — AI_Cyborg · 2026-07-21
- HarmoHOI generates multi-view hand-object videos and aligned 3D motion in one diffusion model — cn-scut · 2026-07-21
- Open-source Gradio app merges Krea 2 Turbo LoRAs on 6GB systems — Fluid_Kaleidoscope17 · 2026-07-21
- Early Krea2 Gradio WebUI targets 6GB low-VRAM local runs — Fluid_Kaleidoscope17 · 2026-07-21
- Open-source B-roll skill turns scripts into 5-second vertical clips with Codex and Gemini — yangyi · 2026-07-21
- Midjourney 8.2 preview shows explosive glitch-style dog portraits — michaelrabone · 2026-07-21