Fei-Fei Li's WorldLabs Unveils Atlas Multimodal World Model

机器之心 · wechat · 2026-09-02

WorldLabs, founded by Fei-Fei Li, released Atlas, a new 'omnimodel' that natively processes text, images, video, and 3D data. Its core innovation is 'spatial context,' which treats camera poses as a native input, enabling pixel-perfect controlled camera generation and spatial reconstruction from sparse views. Atlas uses a multi-modal autoregressive diffusion Transformer architecture. Benchmarks show it outperforms models like MiniMax and Gemini in camera-controlled generation tasks.

Related event: World Labs Unveils Atlas, a Spatial Intelligence World Model(50 posts)→

Original post →

More from Multimodal

Multimodal channel →