SceneAgent: agentic pipeline turns 3D captures into physics-ready scenes for robot training

hankyang94 · x · 2026-09-18

SceneAgent is an agentic pipeline that turns 3D captures into simulatable scenes for robotics policy training and evaluation, combining semantic features with predictive per-Gaussian physics to make individual objects or entire scenes interactive.

The pipeline automates: 3DGS processing from images, video, LiDAR or generated scenes (e.g. World Labs Marble) with per-image calibration at 30k/60k steps; per-Gaussian semantic feature inference; object segmentation and background infill; baking physics materials (rigidity, friction, density); decomposing objects into parts; articulating joints; and generating similar meshes with varied geometry.

Original post →

More from Embodied

Embodied channel →