Robot eval dashboard tracks the 95% lower confidence bound
DominiqueCAPaul · x · 2026-07-20
The author showed a new dashboard for a robot actuator unboxing task.
- The main metric is success rate, but instead of the raw observed rate they track the 95% lower confidence bound.
- The reasoning is to avoid p-hacking from running too many experiments; if they want to improve the number, they can simply run more evals.
- They note the chart was inspired by autoresearch-style dashboards, but real robotics evals are far less automated because every test still needs supervision.
More from Embodied
- Orlando robotaxi ride goes unsupervised in a Model Y — aelluswamy · 2026-07-21
- Anthropic and Physical Intelligence held early M&A talks this spring — steph_palazzolo · 2026-07-21
- EPO improves 3D foundation models by aligning edges, poses, and depth — ducha_aiki · 2026-07-21
- Applied Intuition launches Dana for physical AI, alongside a16z interview on one billion machines — a16z · 2026-07-21
- Hearing Aids Embrace AI: Auracast and OmegaAI Reshape Personal Audio Broadcasting — BrandonSawalich · 2026-07-21
- a16z Talks with Applied Intuition: The Next Decade of Physical AI — a16z Podcast · 2026-07-21