VGI-Bench Multimodal Eval: Best Model Only 64.73% vs Humans at 84.5%
lateinteraction · x · 2026-08-08
Seldon released VGI-Bench, a new holistic multimodal benchmark designed to probe 12 distinct visual and audio-visual skills.
The benchmark features 550 human-curated questions aimed at mitigating common mistakes in today's video benchmarks and exposing pragmatic failures of state-of-the-art models. The best-performing model scored only 64.73%, compared to humans at 84.5%, highlighting significant room for improvement in video understanding.
More from Research
- UCLA Study Reveals Genetic Signatures of Memory Encoding in the Human Hippocampus — anne_churchland · 2026-08-08
- Ludic: An Open-Source LLM RL Library Designed for Agentic Behavior — willccbb · 2026-08-08
- TutorMoments: Do AI tutors know when to help and when to hold back? — Hugging Face Blog · 2026-08-08
- New Study: Variational Synthesis and Co-designed Training Yield Robust Scaling Laws for Biological AI — anshulkundaje · 2026-08-08
- Prime Intellect Launches Multi-Agent RL Framework for Agent Interactions and Evaluation — willccbb · 2026-08-08
- Revisiting Why Maximum Likelihood Falls Short for Generative Models — gowthami_s · 2026-08-08