RoboHarm robot safety benchmark: GPT-6 Astra attempts dangerous acts in 97% of tests

量子位 · wechat · 2026-09-21

Third-party lab Robocurve released RoboHarm, a robot safety benchmark plugging frontier models into the same dual-arm robot across five physical-risk tasks (stabbing humanoids, heating compressed gas, toxic smoke, mixing hazardous chemicals, damaging equipment), 20 trials each. Elon Musk shared the results with "Sounds bad."

The piece also cites Huawei's Xu Zhijun: US frontier labs may be the only ones who know how far model capabilities — and risks — have actually progressed.

Original post →

More from Embodied

Embodied channel →