StreamMAE: self-supervised learning from continuous video streams scales to 95 hours at NeurIPS 2026
y_m_asano · x · 2026-10-09
A NeurIPS 2026 paper by Ivan Martinović, Lukas Knobel and Yuki M. Asano studies self-supervised learning from continuous video streams — frames consumed in temporal order with strict sliding-window batches, no global shuffling or multi-epoch replay. They build WT++, a 95-hour urban walking-tour dataset for streaming pretraining, and find contrastive/distillation methods struggle while MAE is more robust; the key bottleneck is high intra-batch similarity (near-duplicate frames), not inter-batch similarity. StreamMAE adds stream-aware regularization and motion-biased crop selection, outperforming streaming baselines, matching i.i.d. MAE, staying competitive with ImageNet-pretrained MAE, and scaling positively from 12 to 95 hours. The authors also report DINO performs poorly in streaming and will release code and WT++.
Related event: StreamMAE Makes Self-Supervised Learning Work on Continuous Video Streams(4 posts)→
More from Research
- Formalized classification of finite simple groups may be near, says Irving — geoffreyirving · 2026-10-10
- Success-Guided Sampling: zero real-world data for dexterous robot hands — milesmacklin · 2026-10-10
- Google DeepMind launches AI co-clinician initiative for 'triadic care' in medicine — alan_karthi · 2026-10-10
- Anthropic's Geoffrey Irving: is the formalized classification of finite simple groups coming next week? — geoffreyirving · 2026-10-10
- Cell-rejuvenating therapy shows early vision improvement in 2 of 3 glaucoma patients — kimmonismus · 2026-10-10
- Terence Tao: The Era of "Big M" Mathematics and "Big P" Physics Has Arrived — CatAstro_Piyush · 2026-10-10