Hyper3D launches WorldGen: one image to a full editable 3D scene

智东西 · wechat · 2026-09-01

Shanghai-based Hyper3D (Yingmou) released WorldGen, a world-generation model that moves 3D generation from single assets to complete scenes. From one scene photo, it auto-detects major objects (2-3s per-object preview, full scene in 2-3 minutes) and produces multiple independent, editable, interactive assets. Its core CAST architecture won a SIGGRAPH 2025 Best Paper award, recovering occluded structure from a single image and inferring contact/support relations between objects; interactive objects use Mesh while backgrounds use 3D Gaussian Splatting, and a SimReady mode estimates physical attributes like mass and friction.

Applications span embodied AI, gaming, film and XR: partnerships with Dijie Robotics and Mouxianfei for robot simulation training, a July collaboration with Unity China, and film workflows combining 3D scenes with video models like Seedance 2.5 for multi-shot consistency. The team notes it currently supports rigid bodies only, physical properties are inferred rather than measured, and industry adoption is still in beta/validation stages.

Related event: Deemos Releases WorldGen: One Image Generates Fully Editable 3D Worlds(8 posts)→

Original post →

More from Embodied

Embodied channel →