World Labs launches Atlas, an omni world model that rebuilds scenes in 3D from 32 images

theworldlabs · x · 2026-09-18

World Labs (Fei-Fei Li) has introduced Atlas, a next-generation omni world model for spatial intelligence, pretrained from scratch to natively operate on text, images, video, and 3D. It's a multimodal autoregressive diffusion transformer: all inputs combine into a shared spatial context, from which it generates what comes next — staying 3D-consistent with everything seen and imagining beyond it. Performance reportedly improves with training compute.

Four capability pillars:

In the demo, 32 input images were enough to train real-time flight through NVIDIA's Voyager headquarters on Blackwell GPUs. Atlas will power future versions of Marble and other products.

Related event: World Labs Unveils Omni World Model Atlas, Rebuilding NVIDIA HQ From 32 Photos(3 posts)→

Original post →

More from Multimodal

Multimodal channel →