GPT-6 Astra unlocks agentic real2sim: vision models as tools to turn room video into interactive sims

ericjang11 · x · 2026-09-28

Eric Jang rounds up recent work where GPT-6 Astra has enabled new results in robotics and inverse graphics.

The highlighted example is agentic real2sim: instead of asking an LLM to build a Blender scene from a room video, vision models are handed to GPT-6 as tools — Pi3X for geometry, SAM3 for segmentation, and GVHMR/Kimodo for human motion capture — turning a single room video into a higher-quality interactive simulation.

Jang also invites robotics researchers working on robotic control or agentic real2sim to DM for free access to open-source models (Kimi K3, Qwen 3.8 Flash Next) to benchmark their agentic capabilities against GPT-6.

Original post →

More from Embodied

Embodied channel →