AgentVidBench: a multi-hop video QA benchmark testing spatial, temporal and causal reasoning

Kangwook_Lee · x · 2026-09-21

Kangwook Lee's team at UW-Madison released AgentVidBench, a multi-hop video QA benchmark for evaluating MLLM agents on spatial, temporal, and causal reasoning (arXiv + Hugging Face). It goes beyond single-turn scene-level QA with step-by-step solution traces for trajectory evaluation. Tests across 12 proprietary and open-source MLLMs show limited single-turn performance, while agentic workflows improve both accuracy and trajectory scores. Fun fact: it grew out of their private benchmark for building automated video-game dev agents.

Related event: AgentVidBench: A Multi-Hop Video QA Benchmark for MLLM Agents(2 posts)→

Original post →

More from Multimodal

Multimodal channel →