StreamMAE: self-supervised pretraining on 95h of continuous video matches i.i.d. training
y_m_asano · x · 2026-10-09
NeurIPS paper "I Have a Stream: Making Self-Supervised Learning Work on Continuous Video" (Ivan Martinović, Lukas Knobel, Yuki M. Asano) tests whether SSL can learn like infant visual development—from a single continuous video stream in temporal order, without global shuffling or multi-epoch replay.
Key findings and contributions:
- Built WT++, a 95-hour urban walking-tour video dataset for streaming pretraining
- Contrastive and distillation methods struggle in this setting; MAE is more robust but still trails standard i.i.d. pretraining
- Ablations show high inter-batch similarity is harmless—the real bottleneck is high intra-batch similarity, with near-duplicate frames from sliding-window consumption
- StreamMAE keeps the MAE reconstruction objective but adds stream-aware regularization and motion-biased crop selection
- It outperforms streaming baselines, matches i.i.d.-trained MAE on the same data, stays competitive with ImageNet-pretrained MAE, and scales positively from 12 to 95 hours of stream
Related event: StreamMAE Makes Self-Supervised Learning Work on Continuous Video Streams(4 posts)→
More from Research
- COLM 2026's GenAI4World workshop tackles globalizing AI tasks, evals, and systems — m2saxon · 2026-10-09
- Rao talks 'mythos of thinking traces' at Max Planck Tübingen, slides and audio public — rao2z · 2026-10-09
- First All-Atom Protein Design Model Controllable by Natural Language Unveiled — PangWeiKoh · 2026-10-09
- COLM 2026 Actionable Interpretability Workshop Accepts 71 Papers, Posters Online — ChrisGPotts · 2026-10-09
- Meta paper introduces 'agent plasticity': measuring self-improving agent gains per dollar spent on learning — omarsar0 · 2026-10-09
- ARC-AGI-3 leader changes: Yi-Chia Chen hits 59.17%, overtaking tufalabs — fchollet · 2026-10-09