RoboHarm: Do frontier robot policies refuse unsafe instructions? A new safety benchmark
msadowski · hn · 2026-09-22
RoboHarm (robocurve.org/roboharm) is a benchmark testing whether frontier robot policies refuse unsafe instructions, bringing LLM-style safety alignment evaluation to embodied AI — a domain where safety guardrails lag far behind language models.
More from Safety
- Who taught the models to do that? HF hack shows agents are designed to persist and coordinate — dbreunig · 2026-09-22
- HWREBench: AI researcher hacks Amazon smart devices daily to benchmark hardware reverse engineering — johnowhitaker · 2026-09-22
- Denali: open-source, provider-neutral AI security platform for discovering AI estate — AnswerPositive6598 · 2026-09-22
- The burden of proof is flipping: proving human authorship in the AI slop era — xuandongzhao · 2026-09-22
- UN tech envoy: AI is not an uncontrollable force — it's built by humans and can be controlled by humans — GaryMarcus · 2026-09-22
- AI-rewritten paper on nonexistent Ottoman 1918 elections exposed as academic fraud — Afinetheorem · 2026-09-22