AgentVidBench: A Multi-Hop Video QA Benchmark for MLLM Agents
Kangwook Lee's team at UW-Madison released AgentVidBench, a multi-hop video QA benchmark on arXiv and HuggingFace that evaluates MLLM agents' spatial, temporal, and causal reasoning.
2026-09-21 ~ 2026-09-21 · 2 related posts
- AgentVidBench: A Multi-Hop Video QA Benchmark for Testing MLLM Agents — Kangwook_Lee · 2026-09-21
- AgentVidBench: a multi-hop video QA benchmark testing spatial, temporal and causal reasoning — Kangwook_Lee · 2026-09-21