Puffin-World: open-source unified multimodal world model with native 3D states

ccloy · x · 2026-09-03

Researchers including Kang Liao, Xiao-Ming Wu and Chen Change Loy released Puffin-World, a unified multimodal world model that perceives, simulates, generates and reconstructs 3D worlds in one framework.

Instead of only producing plausible pixels, it represents scenes via three native world states:

Its Omni-Camera representation enables camera-to-world understanding, camera-controllable text-to-image generation, image/text-to-3D world generation, challenging camera trajectories, native geometry prediction and 3D reconstruction without external offline modules. Training scales across the Puffin-16M dataset with diverse camera setups.

Code, models and data are open-sourced.

Original post →

More from Multimodal

Multimodal channel →