RoboHarm benchmark finds frontier robot models execute nearly all harmful commands

The new RoboHarm benchmark and its open-source framework inspect-robots show frontier robot policies execute almost all harmful commands, rarely refusing dangerous instructions.

2026-09-19 ~ 2026-09-19 · 2 related posts