TT-VidT: temporal-decoupled video pretraining accepted at NeurIPS
CMHungSteven · x · 2026-09-29
KBlueleaf (Spellbrush) announced that TT-VidT, a video pretraining method, has been accepted at NeurIPS. The approach decouples temporal information to extract real motion cues from video, achieving high-quality results with only small-scale training. Collaborators include NVIDIA and Spellbrush.
More from Research
- ColNanoVDR distills multi-vector visual document retrieval without documents, keeping 95% NDCG@5 at 149M params — nanovdr · 2026-09-29
- NUS rethinks DiT residual connectivity: 1.73x fewer training iterations, 1.39 FID — NationalUniversityofSingapore · 2026-09-29
- Imprint Reader Decodes Weight Updates into Natural Language, Enables Targeted Edits — Guanxu Chen · 2026-09-29
- SJTU's GeoVerse Synthesizes World-Consistent Novel Views in Geometric Latent Space — SJTU · 2026-09-29
- Tencent Hunyuan Maps Scaling Laws for Encoder-Free Multimodal Pretraining — Tencent-Hunyuan · 2026-09-29
- When Do Model Internals Help? Benchmarking Representation Engineering for LLM Safety — Tianyi Guan · 2026-09-29