Closed-loop spatial understanding from just two wrist cameras, no gripper feedback

ChongZzZhang · x · 2026-09-21

A 4x/12x-speed demo shows a robot achieving closed-loop spatial operation using only plain visual understanding from two cameras on the same wrist — no base/world frame observation, no gripper feedback.

Author's takeaways: closed-loop spatial understanding is off-the-shelf, and reasoning and actionability can be coupled without symbolising.

Related event: GPT6 as Policy: Wrist-Mounted Dual Cameras Drive Zero-Shot Robot Stacking(5 posts)→

Original post →

More from Embodied

Embodied channel →