5,500 hours of driving video, a 9B diffusion model: Valeo derives scaling laws for video generation

kwangmoo_yi · x · 2026-09-01

Valeo.ai releases VATIX: using 5,500 hours of real front-facing driving video from the Natix dataset (28 countries, 6.3M 2.5s clips, anonymized), they train video diffusion models from scratch and systematically analyze scaling laws.

Core question: video generation for autonomous driving can't follow web-scale LLM recipes — driving data is expensive, privacy-constrained and limited. Given a fixed dataset, where should compute go?

Method: 200+ training runs from 1.6M to 1.1B parameters fit three scaling laws — model size, training exposure, and optimal allocation under a fixed compute budget, with L(x)=L₀ + A·x^(−α).

Result: extrapolating the fitted laws to a 9B model yields only 3.6% error. Paper, code and dataset are open-source; models are yet to be released. The project page shows side-by-side ground-truth vs generated driving clips.

Related event: Valeo Derives Video Diffusion Scaling Laws from 5,500 Hours of Driving Data(3 posts)→

Original post →

More from Multimodal

Multimodal channel →