StreamMAE: self-supervised learning from continuous video streams scales to 95 hours at NeurIPS 2026

y_m_asano · x · 2026-10-09

A NeurIPS 2026 paper by Ivan Martinović, Lukas Knobel and Yuki M. Asano studies self-supervised learning from continuous video streams — frames consumed in temporal order with strict sliding-window batches, no global shuffling or multi-epoch replay. They build WT++, a 95-hour urban walking-tour dataset for streaming pretraining, and find contrastive/distillation methods struggle while MAE is more robust; the key bottleneck is high intra-batch similarity (near-duplicate frames), not inter-batch similarity. StreamMAE adds stream-aware regularization and motion-biased crop selection, outperforming streaming baselines, matching i.i.d. MAE, staying competitive with ImageNet-pretrained MAE, and scaling positively from 12 to 95 hours. The authors also report DINO performs poorly in streaming and will release code and WT++.

Related event: StreamMAE Makes Self-Supervised Learning Work on Continuous Video Streams(4 posts)→

Original post →

More from Research

Research channel →