RoboHarm benchmark and inspect-robots open-source eval framework for physical AI released

chooi_jeq · x · 2026-09-19

The author released the RoboHarm benchmark with 300 traces, plus inspect-robots, an open-source eval framework for physical AI (544 stars on GitHub). Think Inspect AI for robotics: define a benchmark once, then run any policy (LLM agent or VLA) on any embodiment — a real arm, humanoid, or simulator — with auditable logs (grader scores, LLM transcripts, full config) and first-class Rerun visualization. The project is in early development with a changing API; pin a version before depending on it.

Related event: RoboHarm benchmark finds frontier robot models execute nearly all harmful commands(2 posts)→

Original post →

More from Embodied

Embodied channel →