Pure-vision robot policy: two wrist RGBD cams + claimed "GPT6" hits near-100% success

ChongZzZhang · x · 2026-09-21

A demo (video sped 4x/12x) shows a robot manipulation policy with no base/world-frame observation and no gripper feedback — just plain visual understanding from two RGBD cameras on the same wrist. The author claims a "GPT6 as policy" directly outputs end-effector frame actions from the two cams plus instructions, with near-100% success given inference tokens and time for retries. Tasks include finding an occluded wooden box and placing it on a yellow sign, and stacking boxes. Note the "GPT6" claim is unverified.

Related event: GPT6 as Policy: Wrist-Mounted Dual Cameras Drive Zero-Shot Robot Stacking(5 posts)→

Original post →

More from Embodied

Embodied channel →