LeVJEPA from LeCun's team: video pretraining matches V-JEPA 2 with up to 20.8x less compute
Cohere_Labs · x · 2026-09-03
Cohere Labs' Open Science Community will host LeVJEPA first author Lukas Kuhn for a technical talk and Q&A on Sept 8. The arXiv paper, co-authored by Yann LeCun and Randall Balestriero, introduces the first video encoder trained under LeJEPA's collapse-free objective — dropping EMA target encoders, stop-gradients, and pixel reconstruction in favor of an encoder + projector regularized by SIGReg with a single hyperparameter. At matched epochs, LeVJEPA matches or beats V-JEPA 2 across ViT-S/B/L with 5.6–20.8x less pretraining compute; at matched total FLOPs it exceeds the strongest video baseline by 7.6 points on ImageNet-1K while staying competitive on motion-centric benchmarks. Uniform random token dropping cuts compute and improves downstream accuracy.
Related event: LeVJEPA Cuts Video Pretraining Compute by Up to 20.8x, Matching V-JEPA 2(3 posts)→
More from Multimodal
- Text-to-image style share: dried salt texture prompt demo — LudovicCreator · 2026-09-23
- RefMod tuning on RTX 5090: 20 reference images burn 18,500 tokens for marginal speed gains — spacemidget75 · 2026-09-23
- Qwen-Audio-3.1 models go live: ASR-Flash priced at ¥0.8/M input tokens — aigclink · 2026-09-23
- Alibaba drops five Qwen-Audio-3.1 voice models, cuts ASR pricing by 95% — aigclink · 2026-09-23
- Four invented 'imaginary tokens' to probe Midjourney v8.2's surreal side — LudovicCreator · 2026-09-23
- New ADetailer fork brings ComfyUI Impact-style cycle refinement to Forge WebUIs — sca285 · 2026-09-23