World Labs unveils Atlas world model: native text/image/video/3D, 1-min 1440p camera-controlled video
thione · x · 2026-09-07
World Labs (Fei-Fei Li) introduced Atlas, a next-generation omni world model pretrained from scratch to natively operate on text, images, video, and 3D. It's a multimodal autoregressive diffusion transformer combining all inputs into a shared spatial context, staying 3D-consistent while generating. Capabilities: pixel-perfect camera-controlled generation up to 1 minute of 1440p video; spatial reconstruction from 1-40 images with explicit 3D output claimed to beat specialized SOTA; space-time simulation enabling Real-to-Sim robotics workflows; text-to-image and 360 panoramas. Performance scales with compute; Atlas will power future versions of Marble.
More from Embodied
- Figure's Index creator network processes 30 min of video per second; $15M paid to data creators — adcock_brett · 2026-09-08
- Sim2real nails it first try: a new film production pipeline via robot simulation — alexcovo_eth · 2026-09-08
- Robotics startup Action Intelligence unveils Continuo, debuts at ECCV Sept 10-12 — thetripathi58 · 2026-09-07
- Embodied AI shifts from leaderboards to deployment: Annu Intelligence cuts robot rollout costs by 80% — 智东西 · 2026-09-07
- Unitree's UnifoLM-X2-1.0 world model runs fully autonomous humanoid fights in real time — Distinct-Question-16 · 2026-09-07
- Waymo officially expands to Berkeley as physical AI startup wave builds in the Bay Area — jfiance · 2026-09-07