LeVJEPA: video pretraining at 5.6-20.8x less compute, matching V-JEPA 2
Cohere · youtube · 2026-09-12
LeVJEPA is the first video encoder trained under LeJEPA's collapse-free objective, dropping EMA targets, stop-gradients and pixel reconstruction for a single encoder-plus-projector with one hyperparameter. It matches or beats V-JEPA 2 across ViT-S/B/L at 5.6-20.8x less pretraining compute, exceeds the strongest video baseline by 7.6 points on ImageNet-1K at matched FLOPs, and supports block-causal attention with no accuracy cost — suggesting video is a viable, often preferable substrate for general visual pretraining.
More from Multimodal
- fal launches H3 Max Camera Controls: navigable 3D scenes in under 3 seconds — isidentical · 2026-09-12
- AI storytelling is at the 'animated photographs' stage of early cinema, says Tolan's Eliot Peper — every · 2026-09-12
- ChatGPT Images 2.5 Flare takes 2nd on Human Creativity Benchmark, near Muse Image — yuwen_lu_ · 2026-09-12
- Testing fal's H3 Max Multi Angle video generation with a Wright Brothers prompt — DavidmComfort · 2026-09-12
- fal Podcast Ep. 4: How AI Is Reshaping the Microdrama Industry in China vs the US — gorkem · 2026-09-12
- Redditor Creates a Fake Movie Trailer Entirely with MiniMax Video Models — ColdComplaint8 · 2026-09-12