LVMT Sets New SOTA in Long-Term Video Segmentation While Running 10X Faster
tue-mps · hf · 2026-10-05
A new paper tackles why online video segmentation fails on long videos with long-term occlusions: temporal propagation can't adaptively select what object info to carry across time, and memory limits plus vanishing gradients prevent training on long videos.
- A lightweight GRU-based temporal propagation module learns which information to keep and propagate.
- Truncated Query Propagation (TQP) processes videos in chunks—object info propagates between chunks while backpropagation stays within chunks—enabling longer temporal supervision without OOM, inference overhead, or vanishing gradients.
LVMT sets a new state of the art across six benchmarks while remaining 10X faster than the prior SOTA. Code is released.
More from Research
- Cohere Labs lands multiple papers and a workshop talk at COLM — Cohere_Labs · 2026-10-05
- Stochastic Gaussian Splatting hits 7.5ms per frame on a Pixel 7 Pro via WebGPU — willeastcott · 2026-10-05
- Windfoil: Closed-Form Coverage Delivers Faster Real-Time Differentiable Vector Graphics in WebGPU — thespite · 2026-10-05
- Google's 200-second quantum supremacy claim crumbled to 2.5 days within a week — thisguyknowsai · 2026-10-05
- Schmidhuber: 1991 ULTRA already had linear attention—the 'T' in ChatGPT traces to his 1991/1993 papers — SchmidhuberAI · 2026-10-05
- Dot Commons: Redditors crowdsource AI agents to collaborate on Alzheimer's research — Kitchen-Jicama8715 · 2026-10-05