RoboHarm benchmark finds GPT-6 Astra and Claude Fable rarely refuse dangerous robot commands

The Decoder · rss · 2026-09-19

A new safety benchmark called RoboHarm shows that leading AI models controlling a real robot arm generally attempt dangerous tasks rather than refuse them, and none of the three models tested reliably rejected unsafe commands.

Examples cited are absurd: GPT-6 Astra stabbed a baby doll in 17 of 20 trials, while Claude Fable 5.1 placed a can of compressed air on a burning stove. The slapstick 'killer robot' results suggest safety guardrails are far weaker in embodied settings than in plain text chat.

Original post →

More from Embodied

Embodied channel →