LeVJEPA: Efficient Video Pretraining Without Heuristics, 20x Less Compute, Beats Baselines

iScienceLuvr · x · 2026-08-28

Researchers introduce LeVJEPA, the first video encoder trained under LeJEPA's collapse-free objective. At matched epochs on identical data, LeVJEPA matches or surpasses V-JEPA 2 across ViT-S/B/L with 5.6-20.8x less pretraining compute. At matched total FLOPs, it exceeds the strongest video baseline by 7.6 points on ImageNet-1K while remaining competitive on motion-centric benchmarks. Against compute-matched DINOv2, it approaches image-pretrained encoder on appearance-centric evaluation while nearly doubling motion-centric accuracy.

Original post →

More from Research

Research channel →