RoboFoundry's System-as-Policy self-evolution lifts GPT-5.5 by 27.8% on EmbodiedBench

Jingsong Liang · hf · 2026-09-29

RoboFoundry is the first embodied agentic framework to treat the supporting system itself as a unified policy, rather than optimizing memory, context, skills, or action interfaces in isolation.

It diagnoses capability gaps in decision-making and memory management, converts execution traces into validated task-specific system updates, and promotes recurring improvements to the general system. Evolution operates on two surfaces: a context system managing active context and persistent file-system memory, and a hierarchical skill system organizing atomic skills, reusable compositions, and failure-conditioned recovery. A shared semantic interface separates embodiment-invariant decisions from embodiment-specific execution, enabling transfer across heterogeneous robots.

Results: SOTA on EmbodiedBench, improving GPT-5.5 by 27.8% and bringing Qwen3.7-Plus near parity (70.3% vs 72.7%); at least 39.0% better than all baselines on RoboMemArena for long-horizon memory; and 243.8%–679.7% over Cap-Agent0 across all LIBERO-PRO perturbations. Real-world deployments show zero-shot transfer and online evolution across robots and tasks.

Original post →

More from Embodied

Embodied channel →