I Have a Stream: Making Self-Supervised Learning Work on Continuous Video
Ivan Martinović, Lukas Knobel, Yuki M. Asano
NeurIPS 2026
cs.CV
2026-10-01
StreamMAE pretrains on ordered video streams with only data-pipeline changes and matches i.i.d. MAE on the same frames, gaining further as the stream grows from 12 to 95 hours.
Mainstream self-supervised pretraining (MAE, DINO, MoCo) quietly assumes the dataset can be globally shuffled, so every batch is a set of unrelated images approximating i.i.d. sampling. Visual experience for a robot, a wearable camera, or an edge device does not look like that: frames arrive in temporal order, neighboring frames are near-duplicates, and storing everything for repeated shuffled epochs is often not an option. The paper asks a clean question. Trained from random initialization, with no global shuffling, no multi-epoch replay, and no long-term replay buffer, can SSL still learn good representations from a continuous stream? To make this testable, the authors extend WalkingTours into WT++, 58 public walking-tour videos totaling 95 hours. The distribution gap is easy to see: pairwise DINOv2 feature similarity over 512 consecutive frames averages 0.95, against 0.18 for randomly sampled frames from the same video.
The first move is diagnosis. Benchmarked under the streaming protocol (ViT-S, 12-hour stream), MoCo v3 scores 68.7 on ImageNet-1K, below the 71.9 of a randomly initialized encoder: near-duplicate frames act as false negatives and break contrastive learning. DINO lands at 72.7, barely above random init. MAE is the most stable, but dense prediction degrades, and the degradation grows with model size: ViT-B loses nearly 10 mIoU on Cityscapes (58.2 vs 67.8 for same-data i.i.d.). A reproduced MemoryStoryboard with a replay buffer only reaches 71.0. MAE becomes the base.
The attribution experiment is the best part of the paper. Streaming differs from i.i.d. in two ways: consecutive batches overlap heavily (inter-batch similarity), and each batch is full of near-duplicates (intra-batch similarity, 0.665 on WT++12h vs 0.004 for an ImageNet batch). The authors pre-shuffle ImageNet-1K once and consume it as a fixed sliding-window stream: inter-batch similarity hits 0.989, almost identical to a real video stream, while intra-batch stays at 0.004. This fake stream trains to exact parity with standard i.i.d. MAE (77.4 vs 77.4 top-1, ViT-S). Inter-batch redundancy is harmless; near-duplicate frames inside a batch are the killer. Gradient analysis points the same way: the closer a streaming method's mean consecutive-batch gradient cosine sits to the near-zero level of i.i.d. MAE (0.04), the better it transfers. Streaming MAE averages -0.11; StreamMAE pulls it back to 0.01.
StreamMAE follows from this. The MAE objective is untouched; every change lives in the input pipeline:
| Method (ViT-S, WT++12h) | IN-1K Acc@1 | Cityscapes mIoU | ADE20K mIoU |
| Random init | 71.9 | 47.7 | 16.7 |
| MoCo v3 (stream) | 68.7 | 52.0 | 19.4 |
| MAE (stream) | 77.1 | 61.3 | 23.9 |
| MAE i.i.d. (same data) | 77.0 | 63.5 | 25.9 |
| StreamMAE | 77.5 | 63.8 | 26.1 |
| StreamMAE (95h stream) | 78.4 | 68.0 | 29.5 |
At 12 hours StreamMAE reaches parity with same-data i.i.d. MAE, and passes it with scale: ViT-B on the 12-hour stream gets 69.0 Cityscapes mIoU against 67.8 for i.i.d.; ViT-B on 95 hours reaches 82.0 / 74.0 / 36.5 on IN-1K / Cityscapes / ADE20K, trading blows with an iteration-matched ImageNet-1K i.i.d. MAE at 81.5 / 72.3 / 35.9. Cumulative ablations move Cityscapes from 60.8 (streaming MAE) to 63.8, with regularization and crop selection contributing most. Frozen-representation attentive probing improves by 6.8 (ViT-S) and 9.8 (ViT-B) points over streaming MAE.
Two scaling results are hard numbers worth remembering. Raising the sampling rate from 3.75 to 15 FPS, which quadruples gradient updates, hurts Cityscapes (65.2 vs 69.0). Repeating the 12-hour stream twice (70.3) underperforms feeding a fresh 25-hour stream (72.5). Gains come from new scenes, not more updates. The method also transfers: on HD-EPIC (indoor kitchen egocentric), CROWD (dashcam), and KrishnaCAM (head-mounted daily life), StreamMAE beats streaming MAE on every task, with segmentation gains of 1.6 to 5.8 mIoU.
For embodied agents and edge devices whose data arrives once and in order, this is the first from-scratch recipe that closes the gap to i.i.d. pretraining, and the price is low: the objective is untouched, everything changes in the data pipeline. The diagnosis is reusable on its own. If your training data is a redundant stream (logs, sensors, live feeds), contrastive and self-distillation objectives are off the table and masked reconstruction is the right starting point; fix intra-batch diversity and ignore inter-batch overlap. To be clear about what this is: an incremental improvement. MAE was already the strongest streaming baseline, every StreamMAE component is a known technique, and the novelty lies in a combination motivated by a controlled attribution experiment.
There is no dedicated limitations section; the list below is half author-acknowledged, half from reading the numbers.