World Labs unveils Atlas world model: native text/image/video/3D, 1-min 1440p camera-controlled video

thione · x · 2026-09-07

World Labs (Fei-Fei Li) introduced Atlas, a next-generation omni world model pretrained from scratch to natively operate on text, images, video, and 3D. It's a multimodal autoregressive diffusion transformer combining all inputs into a shared spatial context, staying 3D-consistent while generating. Capabilities: pixel-perfect camera-controlled generation up to 1 minute of 1440p video; spatial reconstruction from 1-40 images with explicit 3D output claimed to beat specialized SOTA; space-time simulation enabling Real-to-Sim robotics workflows; text-to-image and 360 panoramas. Performance scales with compute; Atlas will power future versions of Marble.

Original post →

More from Embodied

Embodied channel →