Fei-Fei Li's World Labs unveils Atlas world model: 3 photos replace 300 for 3D capture
theworldlabs · x · 2026-09-05
World Labs, co-founded by Fei-Fei Li, launched Atlas, a multimodal world model that generates image/video frames with pixel-perfect camera control and reconstructs them in 3D.
- Core idea: where LLMs do next-token and video models next-frame prediction, Atlas is built on new view prediction — the first model to unify pixel generation and reconstruction, two problems CV has split for half a century
- Practical payoff: 50-100x cheaper 3D capture — a room that used to need 100-300 photos now takes just three
- The a16z conversation also covers Matrix-style bullet time done with three iPhones, a one-night Slack message that made them bet the company, and why robotics is data-bottlenecked, not chip-bottlenecked
More from Multimodal
- 360° Scenes From a Single Image: Ben Mildenhall Demos Leap in Novel View Synthesis — ricklamers · 2026-09-05
- Orbis 1.0 tops real-time interactive video benchmarks in quality, physics, and human preference — _akhaliq · 2026-09-05
- Google rolls out Lyria 3.5 music model in Gemini with vocals and longer tracks — cedric_chee · 2026-09-05
- GPT-6 Astra rebuilds Arc du Carrousel in Blender from 1800s schematics — OfirPress · 2026-09-05
- Microsoft's MAI-Image-2.6-Flash takes #3 on image editing leaderboard at $19.5/1k images — ArtificialAnlys · 2026-09-05
- Microsoft's MAI-Image-2.6-Flash hits #3 in image editing, jumping 34 Elo over last Flash — ArtificialAnlys · 2026-09-05