I Have a Stream: Making Self-Supervised Learning Work on Continuous Video
Standard self-supervised learning pipelines sample images independently and shuffle them globally across epochs, differing from infant visual development or continuous streaming. The authors study self-supervised learning from continuous video streams using strict sliding-window batches without global reshuffling. They construct WT++, a 95-hour urban walking-tour video dataset for streaming pretraining. Evaluating contrastive, distillation, and MAE methods, they identify high intra-batch similarity as the core challenge. To mitigate this, they propose StreamMAE, which adapts the input pipeline with stream-aware regularization and motion-biased crop selection. StreamMAE outperforms streaming baselines, matches i.i.d. MAE, and scales positively as the stream grows from 12 to 95 hours.
WT++ is a 95-hour urban walking-tour video dataset constructed for streaming pretraining.
StreamMAE adapts the input pipeline with stream-aware regularization and motion-biased crop selection.
It scales positively as the pretraining stream grows from 12 to 95 hours.