AgentVidBench: A Multi-Hop Video QA Benchmark for MLLM Agents

Kangwook Lee's team at UW-Madison released AgentVidBench, a multi-hop video QA benchmark on arXiv and HuggingFace that evaluates MLLM agents' spatial, temporal, and causal reasoning.

2026-09-21 ~ 2026-09-21 · 2 related posts