AgentVidBench: a multi-hop video QA benchmark testing spatial, temporal and causal reasoning
Kangwook_Lee · x · 2026-09-21
Kangwook Lee's team at UW-Madison released AgentVidBench, a multi-hop video QA benchmark for evaluating MLLM agents on spatial, temporal, and causal reasoning (arXiv + Hugging Face). It goes beyond single-turn scene-level QA with step-by-step solution traces for trajectory evaluation. Tests across 12 proprietary and open-source MLLMs show limited single-turn performance, while agentic workflows improve both accuracy and trajectory scores. Fun fact: it grew out of their private benchmark for building automated video-game dev agents.
Related event: AgentVidBench: A Multi-Hop Video QA Benchmark for MLLM Agents(2 posts)→
More from Multimodal
- Seedance 2.5 AI Video Demo Goes Viral for Striking Quality — SimplyAnnisa · 2026-09-21
- H3 Minimax swap update demo shows off new character-swap results — No_Stomach1386 · 2026-09-21
- AI cover band nails Avenged Sevenfold style — can you even tell? — Lucidjordan79 · 2026-09-21
- Jev Sparse Attention Cuts MiniMax H3 Video Generation Time by 40% — Repulsive_Gap_1678 · 2026-09-21
- Uncensored Qwen-Image-2.1 GGUF quantization trends on Hugging Face — abenzerps · 2026-09-21
- Custom CUDA shim runs Stable Diffusion on Mac faster than RTX 5090 on Windows, up to 61% quicker — LioDavinchy · 2026-09-21