LeVJEPA: collapse-free video pretraining matches V-JEPA 2 with 5.6–20.8x less compute
Cohere_Labs · x · 2026-09-03
Cohere Labs' Open Science Community announced two sessions next week, headlined by LeVJEPA: the first video encoder trained under LeJEPA's collapse-free objective, dropping EMA target encoders, stop-gradients, and masked pixel reconstruction. Key points:
- Invariance loss over global/local views, regularized by SIGReg, with a provable anti-collapse guarantee; architecture reduces to an encoder plus projector and a single hyperparameter.
- Uniform random token dropping cuts pretraining cost while improving downstream accuracy.
- At matched epochs, LeVJEPA matches or beats V-JEPA 2 across ViT-S/B/L with 5.6–20.8x less pretraining compute.
- At matched total FLOPs it exceeds the strongest video baseline by 7.6 points on ImageNet-1K while staying competitive on motion-centric benchmarks.
The second session covers predictive representations and memory for continual RL with Raymond Chua.
Related event: LeVJEPA Cuts Video Pretraining Compute by Up to 20.8x, Matching V-JEPA 2(3 posts)→
More from Research
- Cohere Labs session probes scaling law reliability and offers a research checklist — Cohere_Labs · 2026-09-23
- SoL-Pi paper: auto-research loops save $4-13 per hour on coding agents — alex_verem · 2026-09-23
- NVIDIA's SoL-Pi lets AI rewrite agent harnesses, cutting tokens 44.7-49% — alex_verem · 2026-09-23
- Slingshot RL framework jailbreaks Qwen2.5-32B at 67% success, transfers zero-shot to Gemini 2.5 Flash — j_foerst · 2026-09-23
- TLAPS-Bench launches in alpha to test whether AI agents can formally prove system correctness — tianyin_xu · 2026-09-23
- ImIR replaces text prompts with image instructions for all-in-one restoration — Süleyman Aslan · 2026-09-23