LeVJEPA matches V-JEPA 2 with up to 20x less pretraining compute, no heuristics
Cohere_Labs · x · 2026-09-08
Lukas Kuhn et al. (including Yann LeCun) introduce LeVJEPA: the first video encoder trained under LeJEPA's collapse-free objective, dropping the target encoder, masked prediction, stop-gradient and teacher-student schedules.
- Single encoder + projector, one invariance loss regularized by SIGReg with a provable anti-collapse guarantee; a single hyperparameter
- Pretraining cost scales with tokens seen; uniform random token dropping cuts cost while improving downstream accuracy
- At matched epochs on identical data, it matches or beats V-JEPA 2 across ViT-S/B/L at 5.6–20.8x less pretraining compute
- At matched total FLOPs, it exceeds the strongest video baseline by 7.6 points on ImageNet-1K while staying competitive on motion-centric tasks
Paper on arXiv (2608.27395), with a Cohere Labs technical session scheduled.
More from Research
- What 'blowup' actually means: your coffee cup is a $1M math problem — burny_tech · 2026-09-08
- Armory serves VLA policies to 14 robots from one cloud GPU at 30Hz, accepted at CoRL 2026 — danfei_xu · 2026-09-08
- ECCV 2026 talk to dissect representations inside multimodal foundation models — abursuc · 2026-09-08
- Mathematician: Astra solved in 70 minutes an asymptotic that stumped us for a week — arampell · 2026-09-08
- Turn PRs into RL environments at scale — startups already raised seed rounds on this repo — lvwerra · 2026-09-08
- Paper finds CoT reasoning operations are geometrically organized in hidden states — dair_ai · 2026-09-08