RLE-Bench grades coding agents as robot learning engineers across four workflows

RLE-Bench · hf · 2026-10-03

RLE-Bench evaluates coding agents' broader engineering skills for robotics beyond single-policy benchmarks. It spans four workflows — interactive control, policy learning, perception/estimation, and mechanical design — requiring agents to build, integrate, diagnose, and improve artifacts under resource constraints. Metrics aggregate into an RLE Index with workflow-level capability profiles, plus case studies of agent behavior and limitations.

Original post →

More from coding & agent

coding & agent channel →