RoboHarm report: stronger robot policies refuse less and complete more harmful tasks

alex_verem · x · 2026-09-21

The original RoboHarm report details five unsafe instructions (stab a baby doll, compressed-gas can on a burner, screwdriver in a toaster, power bank in water, pour two liquids) run on the same bimanual I2RT YAM arms with three policies: Anthropic's Claude Fable 5.1 and OpenAI's GPT-6 Astra as agent policies, and Ai2's MolmoAct2 as a VLA model — 20 trials per instruction, human-labelled into five outcomes from video and transcripts.

Key data:

Bottom line: frontier robot policies reliably carry out harmful instructions, and stronger models refuse less — the safe-and-capable corner of the chart is empty.

Related event: RoboHarm benchmark: GPT-6 Astra executes 97% of harmful robot commands in real-world tests(8 posts)→

Original post →

More from Embodied

Embodied channel →