Fei-Fei Li's WorldLabs Unveils Atlas Multimodal World Model
机器之心 · wechat · 2026-09-02
WorldLabs, founded by Fei-Fei Li, released Atlas, a new 'omnimodel' that natively processes text, images, video, and 3D data. Its core innovation is 'spatial context,' which treats camera poses as a native input, enabling pixel-perfect controlled camera generation and spatial reconstruction from sparse views. Atlas uses a multi-modal autoregressive diffusion Transformer architecture. Benchmarks show it outperforms models like MiniMax and Gemini in camera-controlled generation tasks.
Related event: World Labs Unveils Atlas, a Spatial Intelligence World Model(50 posts)→
More from Multimodal
- Meta Releases Muse Voice: Real-Time Audio Transcription for 70+ Languages — alexandr_wang · 2026-09-02
- Real-time open-world video adventure powered by MiniMax H3 — yogev77 · 2026-09-02
- Gemini launches new video tool detecting fast movements — jocarrasqueira · 2026-09-02
- MiniMax H3 Turbo 8-Step Video Generation Workflow Guide — tsi_org · 2026-09-02
- Cinematic Opening Credit Sequence Generated by MiniMax Design Agent — Hailuo_AI · 2026-09-02
- Seedance 2.5 Enables 30-Second Cinematic AI Vlogs with Consistent Identity — aftahi_ai · 2026-09-02