WorldSculpt: Compositional Mesh Reconstruction of Cluttered Scenes from Grounded Video

Muyao Niu · hf · 2026-09-07

WorldSculpt adapts a single-object 3D generative prior to multi-view observations, enabling scalable compositional mesh reconstruction of densely cluttered scenes with severe occlusion.

Instead of holistic reconstruction, the method decomposes a scene into individual objects generated independently and composed into a full world, driven by a grounded video. Useful for 3D content creation and simulation scene building.

Original post →

More from Multimodal

Multimodal channel →