TimeLens2 improves video temporal grounding with set-based supervision

MCG-NJU · hf · 2026-07-21

**TimeLens2** proposes a new approach to generalist video temporal grounding for multimodal LLMs. - The paper targets a harder setting where a model must predict a variable number of evidence intervals across different video lengths, domains, query forms, and viewpoints. - It argues that existing training is misaligned with this set-valued task, especially for long videos and reinforcement-learning rewards. - To fix this, TimeLens2 treats evidence as an interval set end to end, and introduces: - **TimeLens2-93K**, a multi-span supervision dataset built with caption-derived proposals, independent localization, cross-agent consensus, semantic verification, and boundary refinement. - A **temporal Wasserstein reward** based on exact 1D Wasserstein distance between interval supports. - A **temporal IoU** term for precise overlap feedback. - Across seven benchmarks, the 2B model beats size-matched baselines everywhere, while the 4B and 8B variants reach state of the art and outperform open-source models with up to 397B parameters. - Compared with Qwen3-VL backbones, the 2B/4B/8B variants improve by **14.2 / 13.0 / 18.1 mIoU** points respectively.

Original post →

More from Multimodal

Multimodal channel →