AgentVidBench: A Multi-Hop Video QA Benchmark for Testing MLLM Agents
Kangwook_Lee · x · 2026-09-21
Kangwook Lee's team released AgentVidBench on arXiv, a multi-hop video question-answering benchmark evaluating spatial, temporal, and causal reasoning in MLLM agents.
Highlights:
- Existing video benchmarks mostly require single-step inference (scene-level queries or global summaries); AgentVidBench targets multi-hop multimodal reasoning.
- Beyond QA pairs, it provides step-by-step solution traces enabling trajectory evaluation — checking whether agents explicitly gather the evidence needed to justify answers.
- Experiments across 12 proprietary and open-source MLLMs show limited single-turn performance, while agentic workflows generally improve both accuracy and trajectory scores.
- Fun fact: it grew out of the team's private benchmark for automated video-game QA agents.
Related event: AgentVidBench: A Multi-Hop Video QA Benchmark for MLLM Agents(2 posts)→
More from Multimodal
- Qwen 2.1 T2I & edit test: Reddit calls it the open-source Nano Banana moment — LongjumpingGur7623 · 2026-09-21
- GPT-Image 2.5 proves surprisingly good at pixel art, animated via Seedance 2.5 — Aiden_Tech_Ai · 2026-09-21
- Qwen-Image 2.1 runs on SGLang: image generation in 18.7s on a single RTX 4090 — Alibaba_Qwen · 2026-09-21
- Qwen Image 2.1 vs GPT Image 2 vs FLUX.2 Klein 9B — Citadel_Employee · 2026-09-21
- NoSpoon Launches Microdrama and Music Video Agents, Will Go Private This Month — Kyrannio · 2026-09-21
- OmniVBench: a 12k-checklist benchmark and 340K-sample dataset for omni reference-to-video generation — Wenxue Li · 2026-09-21