Parallel Tube Decoding Enables Efficient Spatio-Temporal Video Grounding
ByteDance · hf · 2026-08-31
Parallel Tube Decoding enables simultaneous spatial and temporal video grounding by removing autoregressive dependencies, drastically cutting latency while improving accuracy.
More from Multimodal
- Breeze-TTS-2 demo now available on Hugging Face — BreezeBlue · 2026-09-01
- 求助:Anima 模型训练正常但推理生成模糊 — a_throwawayorsmthn · 2026-09-01
- 求测:Ideogram 4 INT8 量化版在 3060 12G 上的表现 — WhyDoiHearBosssMusic · 2026-09-01
- Help: Swapping a Mustache Using a Reference Image in Flux Workflows — sadboi2021 · 2026-09-01
- Fal's new model heralds next chapter for Hollywood VFX — briannekimmel · 2026-09-01
- Sea Angels Generated in p5.js Using Claude Opus 5 — anselm · 2026-09-01