LeVJEPA: video encoder matches V-JEPA 2 with 5.6-20.8x less pretraining compute
CSProfKGD · x · 2026-09-03
Lukas Kuhn will present LeVJEPA at the Cohere Labs Open Science Community on Sept 8 — the first video encoder trained under LeJEPA's collapse-free objective.
- Prevailing video self-supervised methods avoid representation collapse via architectural asymmetries (EMA target encoder, stop-gradient, capacity-limited predictor) or masked pixel reconstruction, both computationally expensive. LeVJEPA drops both: a single encoder plus projector, an invariance loss over global and local clip views, and SIGReg regularizing collapse away with a provable guarantee — the objective has a single hyperparameter.
- Uniform random token dropping cuts the tokens the encoder sees while improving downstream accuracy.
- At matched epochs, LeVJEPA matches or beats V-JEPA 2 across ViT-S/B/L with 5.6–20.8x less pretraining compute; at matched FLOPs it beats the strongest video baseline by 7.6 points on ImageNet-1K while staying competitive on motion-centric benchmarks.
The online talk is open to all with registration.
More from Research
- Beyond looped transformers: dynamic test-time computation and self-introspection as the real lever — PlisSergey · 2026-09-03
- DeepLoop paper makes looped transformers scalable; rumor claims frontier models are 48 layers looped twice — StartupYou · 2026-09-03
- KAIST's Declarative Attention lets LLMs skip most KV cache reads — kaist-ai · 2026-09-03
- NVIDIA post-training pipeline hits gold-medal IOI performance, topping top humans — nvidia · 2026-09-03
- Kirin builds large-scale animal motion dataset from in-the-wild video for 3D animation — Brian Nlong Zhao · 2026-09-03
- BPCO paper distills a stable PPO recipe for LLM RL, beating GRPO across scales — max_paperclips · 2026-09-03