GPT-6 Astra crushes HumanCLAW embodied benchmark, interaction success nearly triples
LINJIEFUN · x · 2026-09-10
Running GPT-6 Astra through the open-source HumanCLAW-Bench harness (Meta/NTU et al.'s benchmark for VLM "action intelligence" — closing the loop from perception to physical action) yields large jumps over the previous best:
- FindSR 64.9%→75.5% (+10.6), NavSR 42.4%→57.1% (+14.7), InteractSR 16.8%→46.6% (+29.8)
- Gains compound: find-to-nav conversion 65%→76%, nav-to-sit 39%→72%
- Solves 147/507 (29%) episodes missed by all nine prior models
- Best-in-class embodied spatial awareness: lowest disturbed-distance (dDtb 0.65 m) while touching a normal number of objects
- Stronger spatial memory on sitting tasks where the target leaves the egocentric view
Takeaway: general-purpose models are rapidly closing the gap on embodied action, not just perception.
Related event: GPT-6 Astra Dominates Embodied Benchmark, Nearly Tripling Success Rate(2 posts)→
More from Embodied
- MicroSLAM tops LaMaria benchmark, turning cheap egocentric video into robot training environments — lukas_m_ziegler · 2026-09-10
- Musk endorses owner who 'retired from driving' after buying a Tesla FSD — elonmusk · 2026-09-10
- Analysts on iPhone Duo: Apple's patience bought vertical integration rivals can't match — BenBajarin · 2026-09-10
- Vision-force fusion robot dressing handles moving arms: 85% arm coverage across 264 real trials — stepjamUK · 2026-09-10
- DIY omnidirectional robot packs LIDAR, thermal camera and hot-swappable battery — arnie_hacker · 2026-09-10
- Dev praises Apple's new health app and Watch HRV focus, says he'll ditch Ōura — TejasKumar_ · 2026-09-10