InterEvolve Evolves Reward Programs at Test Time to Teach Humanoid Robots New Skills
UIUC-CS · hf · 2026-10-02
UIUC researchers present InterEvolve, a test-time evolution framework for humanoid loco-manipulation that solves never-trained tasks without retraining.
- Key insight: a broad locomotion controller already holds most of the competence a new task needs; it just needs an expressive, executable interface between planning and control
- Two components: an object-aware forward-backward behavioral foundation model that turns new rewards into behavior via object residuals on a frozen body prior, plus tasks specified as staged reward programs that an LLM agent revises in context using execution feedback and a library of verified programs, with a numerical optimizer tuning constants
- Candidates are verified across parallel simulation scenarios, and evolved programs sometimes discover strategies human-designed rewards never tap
- Works on diverse tasks and long-horizon compositions in simulation; evolved skills run autonomously on a physical Unitree G1 from egocentric onboard perception
More from Embodied
- VirtualMacOniPad virtualizes full macOS on jailbroken iPads, running Xcode and Final Cut Pro — tom_doerr · 2026-10-02
- HIDE Benchmark Exposes Memory Gaps in Robotic Manipulation Under Partial Observability — Yansong Shi · 2026-10-02
- Human beats Sharpa humanoid robot in IROS reflex drop-sticks challenge, video shows — TinfoilTricorn · 2026-10-02
- Robotic gripper now runs on Opus-written code, aligning parts via built-in light imaging — ihorbeaver · 2026-10-02
- AMD's World Labs Deal Validates 4D World Models; Chinese Rival Lands $28M Factory Order — 机器之心 · 2026-10-02
- Boston Dynamics' new Atlas hand GR3 drops the pinky: 13-DoF design trade-offs explained — xiaohu · 2026-10-02