World Labs unveils Atlas, an omni world model spanning text, images, video and 3D
gowthami_s · x · 2026-09-02
World Labs introduces Atlas, a next-generation omni world model pretrained from scratch to natively handle text, images, video, and 3D. It is a multimodal autoregressive diffusion transformer that merges all inputs into a shared spatial context and stays 3D-consistent when generating; performance scales with training compute.
Key capabilities:
- Camera-controlled generation: images/videos from one or more inputs with pixel-perfect camera control, up to 1 minute of 1440p video;
- Spatial reconstruction: rebuilds real scenes from 1 to dozens of photos, outputting novel views and explicit 3D, beating specialized 3D-reconstruction SOTA;
- Space-time simulation: reframes videos for dramatic effects and enables Real-to-Sim robotics workflows;
- Image generation: text-to-image and 360 panoramas with complex prompt following, text rendering, and diverse styles.
Atlas will power future versions of Marble and other products; a waitlist is open. Early testers report stunning image and pano generation.
Related event: World Labs Unveils Atlas, a Pixel-Perfect Multimodal World Model(20 posts)→
More from Multimodal
- AI-Generated Drink Ad Features Eye Reflections and Splashes — anthara_ai · 2026-09-02
- Prompt for Fabric and Light Title Sequence via MiniMax H3 — umesh_ai · 2026-09-02
- Interest shifts to Meta's real-time voice transcription model — IndraVahan · 2026-09-02
- Fable 5.1 Generates Cinematic Walkthrough via Code — alexalbert__ · 2026-09-02
- Fei-Fei Li on World Models: A Problem Fundamentally Different from LLMs — drfeifei · 2026-09-02
- World Labs Demonstrates Atlas Connecting World Models to Robotics — drfeifei · 2026-09-02