World Labs launches Atlas: an omni world model natively spanning text, images, video and 3D
YunzhuLiYZ · x · 2026-09-16
Fei-Fei Li's World Labs introduced Atlas, a next-generation omni world model pretrained from scratch to natively operate on text, images, video, and 3D. It's a multimodal autoregressive diffusion transformer that combines all inputs into a shared spatial context and stays 3D-consistent while imagining beyond what it sees, with performance scaling with compute.
Key capabilities:
- Camera-controlled generation: video up to 1 minute at 1440p from one or more images, with pixel-perfect camera control
- Spatial reconstruction: rebuilds real scenes from 1–dozens of images, outputting novel-view frames plus explicit 3D, beating specialized SOTA 3D reconstruction models
- Space-time simulation: models space and time from video, enabling reframing and Real-to-Sim workflows for robotics
- Image generation: text-to-image and 360 panoramas with complex prompt following and text rendering
Atlas will power future versions of Marble. World Labs is also hiring in robot learning (Yunzhu Li, Fei-Fei Li) around Atlas for Robotics and Real-to-Sim-to-Real.
Related event: Fei-Fei Li's World Labs Raises $1.2B and Launches Atlas World Model(2 posts)→
More from Embodied
- Travis Kalanick: Tesla is 'the Google of this era' in the physical AI age — rohanpaul_ai · 2026-09-16
- Slovenia's Deputy PM tries Tesla FSD on public roads: 'doesn't get tired, doesn't fall asleep' — elonmusk · 2026-09-16
- DRS-VPT: feed-forward camera pose estimation from a point cloud scan and a single image — kwangmoo_yi · 2026-09-16
- Bionic Robobird Demonstrates Nature-Mimicking Flapping-Wing Flight — TinfoilTricorn · 2026-09-16
- Robotics researcher pushes back on 'omni embodiment' hype: it's the hands, not the abstraction — chris_j_paxton · 2026-09-16
- StarVLA's VLAct trains VLAs on 16 GPUs by reshaping action representations, not data scaling — jiqizhixin · 2026-09-16