SemVID (ECCV 2026): Token Pruning for Video Temporal Grounding via Evidence Chains

量子位 · wechat · 2026-07-30

Video Temporal Grounding (VTG) requires models to accurately determine event start and end times, not just recognize visual content. SemVID, accepted by ECCV 2026, argues that visual token pruning for VTG cannot merely retain "salient frames" like traditional VideoQA; it must preserve an "evidence chain" that supports temporal localization.

SemVID introduces a training-free token pruning framework. It allocates token budgets based on query relevance and inter-frame changes, explicitly selecting three semantic token types within frames: Object (object evidence), Motion (action transitions), and Context (scene anchors). Experiments show that even at extremely low retention rates, this method stably maintains high boundary localization accuracy, avoiding evidence fragmentation.

Original post →

More from Multimodal

Multimodal channel →