Figure's Helix 2.5 hits 56% zero-shot success across 30 unseen homes, vs 9% baseline

APPSO · wechat · 2026-09-18

Figure released Helix 2.5 and ran a zero-shot, blind evaluation in 30 real Bay Area homes it had never seen: 237 of 420 trials succeeded (56%), versus 9% for the same policy without Index pretraining. Robots folded towels, made beds and picked up objects with no on-site data collection or fine-tuning, with self-correction and retries on long tasks.

Key findings

Index dataset

The Index data app launched August 25: 264K downloads, 108 countries, 44K weekly active creators, 16M+ videos uploaded (peaking at 35 minutes of new video per second, 4.9 years of human work daily). Figure has paid creators $15M and plans $1B+ over the next 12 months to scale data 100x, with five processing stages (screening, anti-cheating, dedup, rebalancing, captioning).

Context and limits

CEO Brett Adcock called it "the most important project we've ever done," with data and compute as the biggest bottlenecks.

Related event: Figure Unveils Helix 2.5: Zero-Shot Household Work in 30 Unseen Homes(20 posts)→

Original post →

More from Embodied

Embodied channel →