Fei-Fei Li's World Labs Unveils World Model Atlas
World Labs, co-founded by Fei-Fei Li, released its multimodal world model Atlas in early September. The three co-founders — Fei-Fei Li, Justin Johnson, and Ben Mildenhall — appeared on the a16z podcast (in conversation with Martin Casado) to explain the model's technical approach and their views on spatial intelligence, and the discussion spread widely.
Confirmed
- Atlas's core capability is novel view prediction: from just a few inputs (e.g., 3 photos), it generates images and video frames from new viewpoints and can reconstruct them in 3D space, with support for pixel-level camera control
- The architectural analogy: LLMs predict the next token, video models predict the next frame, and Atlas extends prediction to new viewpoints / the spatial dimension
- Demos included recreating the "bullet time" effect with three iPhones; the team claims traditional 3D capture costs can be cut by 50-100x
- Multiple sharers (ChrisArmstrong, drfeifei, etc.) relayed the same interview content, with consistent information: 3 photos are enough to reconstruct an entire room
Why it matters
- Novel view prediction is seen as potentially the next modeling paradigm after token prediction, a key bet on the spatial intelligence direction
- If 3D scene reconstruction costs truly drop 50-100x as claimed, it would dramatically lower content production barriers in gaming, film, robotics simulation, and other fields
2026-09-04 ~ 2026-09-05 · 5 related posts
Primary sources
- [source] Fei-Fei Li's World Labs explains Atlas: new view prediction may be the next token prediction — a16z · 2026-09-04
- Fei-Fei Li's World Labs unveils Atlas world model: 3 photos replace 300 for 3D capture — theworldlabs · 2026-09-05
3 near-duplicate retellings: Chris_Armstrong · drfeifei · drfeifei