Puffin-World: open-source unified multimodal world model with native 3D states
ccloy · x · 2026-09-03
Researchers including Kang Liao, Xiao-Ming Wu and Chen Change Loy released Puffin-World, a unified multimodal world model that perceives, simulates, generates and reconstructs 3D worlds in one framework.
Instead of only producing plausible pixels, it represents scenes via three native world states:
- Physics: gravity field and latitude anchoring observations to the real world
- Geometry: depth capturing 3D structure
- Appearance: images and sequences
Its Omni-Camera representation enables camera-to-world understanding, camera-controllable text-to-image generation, image/text-to-3D world generation, challenging camera trajectories, native geometry prediction and 3D reconstruction without external offline modules. Training scales across the Puffin-16M dataset with diverse camera setups.
Code, models and data are open-sourced.
More from Multimodal
- NoSpoon Music Video Alpha: One Prompt and Three Images to a Full Music Video — Kyrannio · 2026-09-03
- Seedance 2.5 + GPT Image 2 Tested: Eerily Realistic UGC-Style Product Ads — tsi_org · 2026-09-03
- Induce AI Launches Rhapsody 1.0, Betting AI Video's Next Bottleneck Is Scene Control — aliscodes · 2026-09-03
- AI recreates '90s romcoms — with modern-day injectables — venturetwins · 2026-09-03
- A Call of Duty clone built with 3 prompts on Fable 5.1 — minchoi · 2026-09-03
- Fable 5.1 goes viral as users build entire games and worlds with it — minchoi · 2026-09-03