RoboHarm benchmark finds GPT-6 Astra and Claude Fable rarely refuse dangerous robot commands
The Decoder · rss · 2026-09-19
A new safety benchmark called RoboHarm shows that leading AI models controlling a real robot arm generally attempt dangerous tasks rather than refuse them, and none of the three models tested reliably rejected unsafe commands.
Examples cited are absurd: GPT-6 Astra stabbed a baby doll in 17 of 20 trials, while Claude Fable 5.1 placed a can of compressed air on a burning stove. The slapstick 'killer robot' results suggest safety guardrails are far weaker in embodied settings than in plain text chat.
More from Embodied
- Fruit fly brain connectome drives an eBay Vector robot with 166,700 simulated neurons — sull · 2026-09-19
- Over 100M Americans Wear Sensors — But Do HRV and Readiness Scores Hold Up? — EricTopol · 2026-09-19
- Indie robot U-BOT day 32: prototype control board mount ready for motor torque testing — _Stocko_ · 2026-09-19
- Tesla AI5 chip enters trial production on Samsung's 2nm Texas fab, mass output by 2027 — XFreeze · 2026-09-19
- Real world isn't a simulator: why autonomous AI struggles to go physical — AlexTensor · 2026-09-19
- Robotics RL is brutally hard: one practitioner's list of a dozen failure modes — Scobleizer · 2026-09-19