TimeLens2 improves video temporal grounding with set-based supervision
MCG-NJU · hf · 2026-07-21
TimeLens2 proposes a new approach to generalist video temporal grounding for multimodal LLMs.
- The paper targets a harder setting where a model must predict a variable number of evidence intervals across different video lengths, domains, query forms, and viewpoints.
- It argues that existing training is misaligned with this set-valued task, especially for long videos and reinforcement-learning rewards.
- To fix this, TimeLens2 treats evidence as an interval set end to end, and introduces:
- TimeLens2-93K, a multi-span supervision dataset built with caption-derived proposals, independent localization, cross-agent consensus, semantic verification, and boundary refinement.
- A temporal Wasserstein reward based on exact 1D Wasserstein distance between interval supports.
- A temporal IoU term for precise overlap feedback.
- Across seven benchmarks, the 2B model beats size-matched baselines everywhere, while the 4B and 8B variants reach state of the art and outperform open-source models with up to 397B parameters.
- Compared with Qwen3-VL backbones, the 2B/4B/8B variants improve by 14.2 / 13.0 / 18.1 mIoU points respectively.
Related event: TimeLens2 Sets New SOTA in Video Temporal Grounding(2 posts)→
More from Multimodal
- Midjourney style code share: --sref 2912175708 — tisch_eins · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- MiniMax H3 MAX nails cooking anime clips: 15-second curry demo with prompts shared — Hailuo_AI · 2026-09-11
- MiniMax Music Production Toolkit 2.5 for ComfyUI adds full mastering chain — Vivid_Promise1700 · 2026-09-11
- New Node Finder for ComfyUI ranks fresh nodes by star velocity and recency — Luke2642 · 2026-09-11
- Using a finisher move on one mosquito with MiniMax H3 MAX — the bug survives — Hailuo_AI · 2026-09-11