LeVJEPA: Preventing video representation collapse with SIGReg
burkov · x · 2026-08-31
Video is a richer training signal than images, but learning from unlabeled video often requires complex methods to prevent "representation collapse." This paper from AmiLab and others investigates if that complexity is necessary.
LeVJEPA trains a single video transformer by aligning representations of different views of the same clip. It uses a statistical regularizer called SIGReg to keep representations distributed, preventing collapse without a second encoder or predictor. A practical consequence is that the model can discard 95% of video tokens during training.
More from Research
- Training RL Policy with Massive Rigid Bodies and Obstacles — yacineMTB · 2026-08-31
- 2011 Paper Reveals Origin of Diffusion Models in Denoising Autoencoders — cloneofsimo · 2026-08-31
- Chinese Team Uses PINNs to Solve Boson Star Families, Overcoming Traditional Numerical Limits — drscotthawley · 2026-08-31
- SenseNova-Vision Formulates Vision as Unified Multimodal Generation — rsasaki0109 · 2026-08-31
- Dietterich: A paper is a structured argument, not a record of how evidence was assembled — tdietterich · 2026-08-31
- AI agents spent $3K on research, papers rejected: failure of judgment — rohanpaul_ai · 2026-08-31