LeVJEPA: video pretraining at 5.6-20.8x less compute, matching V-JEPA 2

Cohere · youtube · 2026-09-12

LeVJEPA is the first video encoder trained under LeJEPA's collapse-free objective, dropping EMA targets, stop-gradients and pixel reconstruction for a single encoder-plus-projector with one hyperparameter. It matches or beats V-JEPA 2 across ViT-S/B/L at 5.6-20.8x less pretraining compute, exceeds the strongest video baseline by 7.6 points on ImageNet-1K at matched FLOPs, and supports block-causal attention with no accuracy cost — suggesting video is a viable, often preferable substrate for general visual pretraining.

Original post →

More from Multimodal

Multimodal channel →