LeVJEPA from LeCun's team: video pretraining matches V-JEPA 2 with up to 20.8x less compute

Cohere_Labs · x · 2026-09-03

Cohere Labs' Open Science Community will host LeVJEPA first author Lukas Kuhn for a technical talk and Q&A on Sept 8. The arXiv paper, co-authored by Yann LeCun and Randall Balestriero, introduces the first video encoder trained under LeJEPA's collapse-free objective — dropping EMA target encoders, stop-gradients, and pixel reconstruction in favor of an encoder + projector regularized by SIGReg with a single hyperparameter. At matched epochs, LeVJEPA matches or beats V-JEPA 2 across ViT-S/B/L with 5.6–20.8x less pretraining compute; at matched total FLOPs it exceeds the strongest video baseline by 7.6 points on ImageNet-1K while staying competitive on motion-centric benchmarks. Uniform random token dropping cuts compute and improves downstream accuracy.

Related event: LeVJEPA Cuts Video Pretraining Compute by Up to 20.8x, Matching V-JEPA 2(3 posts)→

Original post →

More from Multimodal

Multimodal channel →