GPT-6 Astra vs MediaPipe on 3D hand pose: 3 min per frame vs 20 ms

chris_j_paxton · x · 2026-09-08

Developer kstonekuan benchmarked GPT-6 Astra on 3D hand pose estimation, reporting SOTA-level visual/spatial reasoning (MazeBench) and potential for labeling robotics datasets. Constrained to MediaPipe's output schema and visualized side by side, the model's high-reasoning mode took 3 minutes per frame versus MediaPipe's 20 ms — a 9000x gap. MediaPipe can't handle gloved hands, which is where a reasoning model could help. Commenter chrisjpaxton notes MediaPipe is old and most serious data teams have better in-house tools, but it remains highly useful.

Original post →

More from Embodied

Embodied channel →