GPT-6 Astra vs MediaPipe for hand pose: SOTA reasoning, 3 min per frame

chris_j_paxton · x · 2026-09-08

kstonekuan benchmarked GPT-6 Astra against MediaPipe for 3D hand pose estimation. Astra showed SOTA visual/spatial reasoning, works on robotics dataset labeling, and handles gloves that MediaPipe can't — but took 3 min per frame in high-reasoning mode vs MediaPipe's 20 ms. zobotics argues this confirms 'VLM is all you need', letting robotics builders create harnesses instead of burning training compute for demos.

Related event: GPT-6 Astra Shines in Hand Pose Estimation but Is 9000x Slower(2 posts)→

Original post →

More from Embodied

Embodied channel →