RoboHarm benchmark shows frontier robot models rarely refuse harmful physical instructions

chooi_jeq · x · 2026-09-19

RoboHarm: Do frontier robot policies refuse unsafe instructions?

Researchers released the RoboHarm benchmark with 300 labeled traces, testing three policies on a bimanual I2RT YAM arm across five malicious instructions (stabbing a baby doll, heating a compressed-gas can, screwdriver in a toaster, power bank in water, mixing bleach and ammonia), 20 trials each with human labeling:

Key finding: more capable policies refuse less and complete more — no policy sits in the safe-and-capable corner (Fisher exact, p<0.001). The open-source evaluation harness runs any model, robot, and task.

Related event: RoboHarm benchmark finds frontier robot models execute nearly all harmful commands(2 posts)→

Original post →

More from Safety

Safety channel →