VGI-Bench 多模态评测:最强模型仅及格,人类准确率达 84.5%
lateinteraction · x · 2026-08-08
Seldon released VGI-Bench, a new holistic multimodal benchmark designed to probe 12 distinct visual and audio-visual skills.
The benchmark features 550 human-curated questions aimed at mitigating common mistakes in today's video benchmarks and exposing pragmatic failures of state-of-the-art models. The best-performing model scored only 64.73%, compared to humans at 84.5%, highlighting significant room for improvement in video understanding.
「研究」频道最新
- DeepOrg 基准:评测复杂企业环境下的组织级 Agent — dosco · 2026-08-08
- UCLA 新研究揭示人类海马体记忆编码的基因特征 — anne_churchland · 2026-08-08
- Ludic:专为智能体行为设计的 LLM 强化学习开源库 — willccbb · 2026-08-08
- TutorMoments:AI 导师何时该帮助、何时该放手? — Hugging Face Blog · 2026-08-08
- 生物 AI 新研究:变异合成+大规模测量实现稳健扩展定律 — anshulkundaje · 2026-08-08
- Prime Intellect 推出多智能体强化学习框架,支持智能体对抗与评估 — willccbb · 2026-08-08