BiGym 2.0: LLM-Written Policies Hit 65% From One Demo, But π0.5 Still Leads at 75%
stepjamUK · x · 2026-10-09
Researchers released BiGym 2.0, a whole-body humanoid loco-manipulation benchmark for household tasks, with the first thorough comparison of learned policies vs LLM-written robot policies.
Key results: GPT-6 Astra lets coding agents write a humanoid controller from just 1 demo, reaching 65% success — surprisingly far, in a regime where learned policies are useless. But π0.5 trained on 60 demos still leads at 75%, especially on precise object rearrangement.
Verdict: learned policies (from scratch or VLAs) remain the more general solution, handling a broader range of tasks at the cost of more demos.
More from Embodied
- IRVL Researcher to Present UHAS, HRT1 and VLA-Replica at Texas A&M Robotics Seminar — YuXiang_IRVL · 2026-10-09
- Google Gemma-Backed Hardware Hackathon at AGI House Gathers 120 Builders for On-Device AI — agihouse_org · 2026-10-09
- Odyssey launches Odyssey-3, claims SOTA world model on Physics-IQ benchmark — Scobleizer · 2026-10-09
- Figure valued at $39B: tracking Brett Adcock's promise of 100,000 robots in 4 years — annatonger · 2026-10-09
- Sentdex: a raw multimodal LLM drives a robot with zero training, VLA era may be over — Sentdex · 2026-10-09
- Dot users complain the iPhone app still lacks CallKit for calls — cheezemink · 2026-10-09