LeVJEPA: Efficient Video Pretraining Without Heuristics, 20x Less Compute, Beats Baselines
iScienceLuvr · x · 2026-08-28
Researchers introduce LeVJEPA, the first video encoder trained under LeJEPA's collapse-free objective. At matched epochs on identical data, LeVJEPA matches or surpasses V-JEPA 2 across ViT-S/B/L with 5.6-20.8x less pretraining compute. At matched total FLOPs, it exceeds the strongest video baseline by 7.6 points on ImageNet-1K while remaining competitive on motion-centric benchmarks. Against compute-matched DINOv2, it approaches image-pretrained encoder on appearance-centric evaluation while nearly doubling motion-centric accuracy.
More from Research
- Decagon: Detecting Relevant Speaker Changes by Combining Speaker Embeddings with Audio-Native Models — Scobleizer · 2026-08-28
- Code World Model: Coding Agent as the World Brain for Video Generation — Scobleizer · 2026-08-28
- Research on Agent Topology Will Enhance Task Optimization Intuition — cephaloform · 2026-08-28
- Questioning the Advantages of Multi-Agent Systems vs. Single Agents — yoavgo · 2026-08-28
- Microduck RL Training Library Open Source: Head Tracking and Backlash Simulation Tricks — Thom_Wolf · 2026-08-28
- AI Engineer Paris Talk: Post-Training and On-Device Agents — helloiamleonie · 2026-08-28