AWSM grounds LLM-agent 3D scene reconstruction in IMU, depth and pose evidence

anselm · x · 2026-10-02

AWSM addresses a real flaw in LLM-agent 3D reconstruction: models like GPT6 and Astra can build editable scenes from visual observations, but the results can be significantly wrong — curved corners become square, passage widths drift, distances are off. For embodied agents these errors change where a robot can actually move. AWSM grounds agentic reconstruction in geometric evidence: IMU-informed scale cues, depth estimates, and camera poses, producing more faithful, editable worlds for mapping and simulation. A multi-robot demo had four robots follow predefined routes to reception and line up. The broader vision: phygital (physical + digital) worlds agents can navigate, remember, interact with, and learn inside.

Original post →

More from Embodied

Embodied channel →