TimeLens2 improves video temporal grounding with set-based supervision
MCG-NJU · hf · 2026-07-21
**TimeLens2** proposes a new approach to generalist video temporal grounding for multimodal LLMs. - The paper targets a harder setting where a model must predict a variable number of evidence intervals across different video lengths, domains, query forms, and viewpoints. - It argues that existing training is misaligned with this set-valued task, especially for long videos and reinforcement-learning rewards. - To fix this, TimeLens2 treats evidence as an interval set end to end, and introduces: - **TimeLens2-93K**, a multi-span supervision dataset built with caption-derived proposals, independent localization, cross-agent consensus, semantic verification, and boundary refinement. - A **temporal Wasserstein reward** based on exact 1D Wasserstein distance between interval supports. - A **temporal IoU** term for precise overlap feedback. - Across seven benchmarks, the 2B model beats size-matched baselines everywhere, while the 4B and 8B variants reach state of the art and outperform open-source models with up to 397B parameters. - Compared with Qwen3-VL backbones, the 2B/4B/8B variants improve by **14.2 / 13.0 / 18.1 mIoU** points respectively.
More from Multimodal
- Same prompt, Seedance 2 and Grok are compared on cinematic transformation output — LudovicCreator · 2026-07-21
- CG Chefs Showcases Retro Anime Style AI Video Generation — nicolascraske · 2026-07-21
- Night-party video demo uses Seedance 2.0, timecode prompts and 4K upscaling — gen_ericai · 2026-07-21
- Fable + Runway turns a simple prompt into superhero bunnies — tlakomy · 2026-07-21
- LTX-2.3 Foley LoRA turns silent video into generated sound effects — linoy_tsaban · 2026-07-21
- A prompt for Magnific and GPT-2 produced a dense surreal comic about unlived lives — CurieuxExplorer · 2026-07-21