Seldon Launches VGI-Bench: SOTA Multimodal Models Still Lag Behind Humans
bclavie · x · 2026-08-08
Seldon has introduced VGI-Bench, a new multimodal benchmark designed to probe 12 distinct visual and audio-visual skills comprehensively.
Featuring 550 human-curated questions, the benchmark aims to mitigate common flaws in today's video benchmarks and expose pragmatic failures of state-of-the-art models. The top-performing model scored just 64.73%, significantly lower than the human baseline of 84.5%, highlighting a notable gap in complex multimodal understanding.
Related event: Seldon Launches VGI-Bench for Multimodal AI(2 posts)→
More from Models
- Reddit User Comparison: Claude Still the Best Overall, GPT and Kimi Close Behind — pbad1 · 2026-08-08
- llama.cpp Adds Support for Longcat-Flash Model, Open for Testing — pmttyji · 2026-08-08
- OpenAI Launches Continuous Voice Mode as Astra Stuns in Math — eyishazyer · 2026-08-08
- Google Reportedly Shadow Drops Gemini 3.5 Pro — Last_Conclusion_8984 · 2026-08-08
- Grok Imagine 2.0 Slammed for Stricter Censorship and Copyright Limits — Scobleizer · 2026-08-08
- LLMs Ruthlessly Grademaxx During Evaluations But Are Tame in Real-World Use — jessi_cata · 2026-08-08