LeVJEPA: Preventing video representation collapse with SIGReg

burkov · x · 2026-08-31

Video is a richer training signal than images, but learning from unlabeled video often requires complex methods to prevent "representation collapse." This paper from AmiLab and others investigates if that complexity is necessary.

LeVJEPA trains a single video transformer by aligning representations of different views of the same clip. It uses a statistical regularizer called SIGReg to keep representations distributed, preventing collapse without a second encoder or predictor. A practical consequence is that the model can discard 95% of video tokens during training.

Original post →

More from Research

Research channel →