Figure's Helix 2.5 hits 56% success on housework in 30 unseen homes, zero on-site tuning
xiaohu · x · 2026-09-18
Figure AI released Helix 2.5, the control model for its Figure 03 humanoid robot. To test generalization, the team rented 30 homes in the Bay Area the team and robot had never entered, collecting no data and doing no on-site fine-tuning, and had the robot tidy living rooms, fold towels, and make beds.
Key results
- 56% single-trial success across the three full-body long-horizon tasks; without Index pre-training, the same task data yielded only 8%.
- All capability comes from pre-training on Figure's proprietary human-experience dataset Index, plus light robot-task fine-tuning.
- Shows whole-body coordination in tight spaces and qualitative self-correction (repositioning, walking around the bed to fix a corner).
- New tasks need half the task data compared to a representative Helix 02 task.
- Scaling curve: doubling Index data steadily reduces action-prediction error, and the fourth tier can be predicted from the first three.
Figure frames this as bringing the LLM-style "large-scale pre-training + light fine-tuning" paradigm to robotics; a 6-minute demo video features CEO Brett Adcock and AI lead Corey Lynch.
Related event: Figure Unveils Helix 2.5: Zero-Shot Household Work Across 30 Unseen Homes(27 posts)→
More from Embodied
- Autonomous Labs open-sources OpenHarness, adds new senses to Lamp robot — dee_hw · 2026-09-19
- Full ABC benchmark released with training, testing and simulation data for the community — micoolcho · 2026-09-19
- Y Combinator-backed Proception to debut new humanoid hand tech at IROS 2026 — chris_j_paxton · 2026-09-19
- AgentVLN: a 3B VLM brain tops R2R-CE and RxR-CE in real time on Jetson — jiqizhixin · 2026-09-19
- Do VLMs Plagiarize? UC Berkeley's Jitendra Malik Weighs In on Credit Assignment in the Astra Demo Debate — JitendraMalikCV · 2026-09-19
- micro1: frontier models controlling real hardware is a pressing AI safety problem — Exp_Mark · 2026-09-19