Seldon Launches VGI-Bench: SOTA Multimodal Models Still Lag Behind Humans

bclavie · x · 2026-08-08

Seldon has introduced VGI-Bench, a new multimodal benchmark designed to probe 12 distinct visual and audio-visual skills comprehensively.

Featuring 550 human-curated questions, the benchmark aims to mitigate common flaws in today's video benchmarks and expose pragmatic failures of state-of-the-art models. The top-performing model scored just 64.73%, significantly lower than the human baseline of 84.5%, highlighting a notable gap in complex multimodal understanding.

Related event: Seldon Launches VGI-Bench for Multimodal AI(2 posts)→

Original post →

More from Models

Models channel →