Dawn Robotics Open-Sources Puffin-World, a World Model Grounded in World State

jiqizhixin · x · 2026-09-20

Dawn Robotics and NTU S-Lab present Puffin-World, a unified multimodal world model natively grounded in world state.

Key idea: for robot training, visually realistic video alone is insufficient — you need environment information supporting perception, localization, planning and action: current camera pose, stable gravity direction, continuous scene geometry, and the next view after movement. Puffin-World goes beyond relative viewpoint changes to include absolute camera orientation relative to gravity and the real world.

The team open-sources Puffin-16M: 15M vision-language-camera triples, 1M challenging rotation trajectories, and absolute camera pose annotations for 44.5M frames.

Original post →

More from Embodied

Embodied channel →