VGI-Bench 多模态评测:最强模型仅及格,人类准确率达 84.5%

lateinteraction · x · 2026-08-08

Seldon released VGI-Bench, a new holistic multimodal benchmark designed to probe 12 distinct visual and audio-visual skills.

The benchmark features 550 human-curated questions aimed at mitigating common mistakes in today's video benchmarks and exposing pragmatic failures of state-of-the-art models. The best-performing model scored only 64.73%, compared to humans at 84.5%, highlighting significant room for improvement in video understanding.

原文链接 →

「研究」频道最新

更多「研究」频道 AI 资讯 →