RoboHarm benchmark and inspect-robots open-source eval framework for physical AI released
chooi_jeq · x · 2026-09-19
The author released the RoboHarm benchmark with 300 traces, plus inspect-robots, an open-source eval framework for physical AI (544 stars on GitHub). Think Inspect AI for robotics: define a benchmark once, then run any policy (LLM agent or VLA) on any embodiment — a real arm, humanoid, or simulator — with auditable logs (grader scores, LLM transcripts, full config) and first-class Rerun visualization. The project is in early development with a changing API; pin a version before depending on it.
More from Embodied
- LEXSUS robots grill live for 550K viewers, self-correcting failures to prove cross-embodiment AI — 创业邦 · 2026-09-20
- Zeno AI shows humanoid robots collaborating on chores with just 4 hours of robot-to-robot data — CyberRobooo · 2026-09-20
- GAVEL harness lifts Qwen3-8B from 41.2% to 91.8% on long-horizon robot tasks — dair_ai · 2026-09-20
- Musk confirms every Starlink V3 satellite will carry an Nvidia Vera Rubin NVL72 — 100k sats, 25GW — ns123abc · 2026-09-20
- Odyssey-3: One Pretrained World Model Adapts to Different Robot Arms With Hours of Demos — ChongZzZhang · 2026-09-20
- iPhone 18 Pro motherboard weighs just 13.1g yet packs PC-class compute — SumitGup · 2026-09-20