TimeLens2 claims SOTA on 7 video grounding benchmarks with 4B and 8B models
_akhaliq · x · 2026-07-21
TimeLens2 is a generalist video temporal grounding MLLM that predicts when supporting evidence appears in videos.
- It reports SOTA across 7 benchmarks.
- The 4B and 8B variants are said to outperform a 397B-parameter model.
- The figure highlights performance across grounding dimensions such as short indoor action grounding, event-level grounding, moments and highlights detection, long-video multi-span grounding, user-style grounding, interrogative grounding, and first-person egocentric grounding.
Related event: TimeLens2 Sets New SOTA in Video Temporal Grounding(2 posts)→
More from Multimodal
- Invideo launches agent-driven video editor that executes edits from plain descriptions — azed_ai · 2026-09-11
- YuE2 music generation gets native ComfyUI support via new PR — LatentSpacer · 2026-09-11
- Mi-Ripple fixes ripple artifacts left by iterative AI image editing — Miyang-AI · 2026-09-11
- Scottish man strolling through his castle: the AI video everyone is sharing — EternalSnow05 · 2026-09-11
- One prompt, full UGC ad: Kling MCP turns a product idea into ready-to-post video — SimplyAnnisa · 2026-09-11
- A Seedance 2.5 quick-start prompt with GPT Image 2.5 hacks — techhalla · 2026-09-11