LeRobot Unifies Robot Evaluation CLI
RemiCadene · x · 2026-07-15
LeRobot 0.6.0 introduces a unified lerobot-eval CLI to evaluate various VLA robot benchmarks, eliminating the hassle of juggling multiple repositories and environments.
This release adds 6 simulation benchmarks, each equipped with dedicated documentation, Docker images, and CI-verified SmolVLA baselines:
- LIBERO-plus: 10,000 perturbation variants across 7 dimensions to observe policy failure conditions
- RoboTwin 2.0: 50 bimanual tasks based on SAPIEN, with 100,000+ trainable trajectories on the Hub
- RoboCasa365: 365 kitchen tasks across 2,500 procedurally generated kitchens
- RoboCerebra: Long-horizon tasks chaining 3–6 sub-goals
- RoboMME: Memory tests including counting, tracking, and imitation
- VLABench: Knowledge and reasoning tasks, from physics to end-to-end "making coffee"
Combined with LIBERO, Meta-World, and NVIDIA IsaacLab-Arena, there are now 9 benchmark families runnable under a single toolset.
More from Embodied
- NVIDIA brings its Cosmos 3 Edge world model to Jetson for on-device robot control — liu_mingyu · 2026-07-21
- A set of agent skills for CAD, robotics, and hardware design — earthtojake · 2026-07-21
- DIY wooden box packs 6 Intel Arc Pro B70 cards with a FreeCAD model — nick_ziv · 2026-07-21
- Snake-like robot moves on fully passive wheels and winding motion — ___Mufasaa · 2026-07-21
- Creator buys a Reachy robot and asks what to build first — dee_hw · 2026-07-21
- MW team shows Gen1 of MW-bot, a semi-humanoid home robot built for pantry storage and ceiling rails — CyberRobooo · 2026-07-21