SemVID (ECCV 2026): Token Pruning for Video Temporal Grounding via Evidence Chains
量子位 · wechat · 2026-07-30
Video Temporal Grounding (VTG) requires models to accurately determine event start and end times, not just recognize visual content. SemVID, accepted by ECCV 2026, argues that visual token pruning for VTG cannot merely retain "salient frames" like traditional VideoQA; it must preserve an "evidence chain" that supports temporal localization.
SemVID introduces a training-free token pruning framework. It allocates token budgets based on query relevance and inter-frame changes, explicitly selecting three semantic token types within frames: Object (object evidence), Motion (action transitions), and Context (scene anchors). Experiments show that even at extremely low retention rates, this method stably maintains high boundary localization accuracy, avoiding evidence fragmentation.
More from Multimodal
- Exploring LoRA Training Settings for Ideogram 4: 1700 Steps with 24 Images — Zealousideal-Car4724 · 2026-07-30
- Audio8-TTS-Preview-0.6b Tops Hugging Face Trending Models — Audio8 · 2026-07-30
- AI Director Creates Full Game Flow Using Kimi K3 and Seedance 2.0 — Xianbao_QIAN · 2026-07-30
- VAST at SIGGRAPH: Generating Animatable 3D Assets in Seconds with 5 Papers — 机器之心 · 2026-07-30
- Flux 3 Text-to-Video Test: Nails French and Auto-Integrates Website Screenshots — OdinLovis · 2026-07-30
- Chinese AI Model Shows Off 'The Greatest Free Kick' — ZabihullahAtal · 2026-07-30