GPT-6 Astra crushes HumanCLAW embodied benchmark, interaction success nearly triples

LINJIEFUN · x · 2026-09-10

Running GPT-6 Astra through the open-source HumanCLAW-Bench harness (Meta/NTU et al.'s benchmark for VLM "action intelligence" — closing the loop from perception to physical action) yields large jumps over the previous best:

Takeaway: general-purpose models are rapidly closing the gap on embodied action, not just perception.

Related event: GPT-6 Astra Dominates Embodied Benchmark, Nearly Tripling Success Rate(2 posts)→

Original post →

More from Embodied

Embodied channel →