LeVJEPA Achieves 20x Efficiency Gain in Video Pretraining

burny_tech · x · 2026-08-30

The LeVJEPA paper introduces a simpler, cheaper video pretraining approach by replacing complex architectural asymmetries with a single encoder trained using invariance and SIGReg regularization. By dropping 95% of video tokens, it matches or beats V-JEPA 2 with 5.6–20.8x less compute. It also exceeds the strongest video baseline by 7.6 points on ImageNet-1K at matched FLOPs. The method allows for block-causal attention without accuracy loss, making temporal ordering a property of the encoder itself.

Related event: LeCun's LeVJEPA Cuts Video Pretraining Compute by Over 20x(4 posts)→

Original post →

More from Research

Research channel →