World Labs releases Atlas, an omni world model native to text, images, video, and 3D
theworldlabs · x · 2026-09-24
World Labs has introduced Atlas, its next-generation omni world model for spatial intelligence, pretrained from scratch to natively operate on text, images, video, and 3D. Architecturally it is a multimodal autoregressive diffusion transformer: all inputs combine into a shared spatial context, and generation stays 3D-consistent with everything seen.
Key capabilities:
- Camera-controlled generation: pixel-perfect camera control from one or more images, up to 1 minute of video at 1440p;
- Spatial reconstruction: reconstructs real scenes from 1 to dozens of images, producing novel-view frames and explicit 3D outputs, outperforming specialized 3D-reconstruction SOTA;
- Space-time simulation: models space and time from video, enabling reframing and Real-to-Sim workflows for robotics;
- Image generation: text-to-image and 360 panoramas with complex prompt following and text rendering.
Atlas scales with training compute, a trend the team expects to hold. It will power future versions of Marble; beta signups are open.
Related event: Atlas Teases Chisel in Beta: Block Out a World and Let AI Bring It to Life(2 posts)→
More from Multimodal
- Google AI Studio rolls out new text-to-speech models with multilingual support — DynamicWebPaige · 2026-09-24
- Claude Opus 5.5 generates a music video in ~1 shot: "not good, but not without interest" — NathanpmYoung · 2026-09-24
- Gemini 3.8 Flash TTS and Flash-Lite TTS land on Merge Gateway, top Hume voice quality index — shensi · 2026-09-24
- LemonSlice Launches Character World Model-1, a Real-Time Interactive Avatar Model — mhdfaran · 2026-09-24
- Reverse workflow: unpack a reference image with Extract Prompt, then remix it — JaynitMakwana · 2026-09-24
- Unverified claim: 'GPT-6 Astra' builds full video projects via Codex + Dreamina CLI — JaynitMakwana · 2026-09-24