Fei-Fei Li's World Labs explains Atlas: new view prediction may be the next token prediction
a16z · youtube · 2026-09-04
On the a16z podcast, World Labs co-founders Fei-Fei Li, Justin Johnson, and Ben Mildenhall discuss Atlas, their latest world model, and the case for spatial intelligence.
Key points:
- Atlas centers on "new view prediction": given images of a scene, the model predicts how that environment should look from a different position in space and time, unifying generation and 3D reconstruction in one model
- The team frames view prediction as a potential primitive for understanding the physical world — analogous to next-token prediction for LLMs
- They cover the technical bets behind Atlas, what it can and can't yet capture, and why dynamics, editability, and simulation matter as world models mature
- The conversation tackles whether Atlas is just a scaled-up video model or a new architecture, video models vs. world models, and whether walkable 4D video is coming
- Applications span creative work, games, architecture, and robotics; Fei-Fei Li argues access to real-world training data is one of robotics' biggest constraints today
Full timestamps and the World Labs blog post on Atlas are linked in the description.
More from Multimodal
- User edits cinematic Tesla Cybercab video entirely with Grok Build — elonmusk · 2026-09-04
- Redditor fine-tunes SDXL on 60 childhood photos to simulate memory recall — uisato · 2026-09-04
- Recreating a viral fight video with MiniMax H3: full reference-generation workflow — MixZealousideal9359 · 2026-09-04
- Midjourney one-word prompting: obscure dialect word 'Sillion' makes a striking image — tisch_eins · 2026-09-04
- Prompt-to-World: Claude Plus Thrixel's Build World Skill Spawns Interactive 3D Worlds — RanaHanocka · 2026-09-04
- Creators showcase AI experiments: synesthesia MIDI, animation galleries, abstract art — floguo · 2026-09-04