BiGym 2.0: LLM-Written Policies Hit 65% From One Demo, But π0.5 Still Leads at 75%

stepjamUK · x · 2026-10-09

Researchers released BiGym 2.0, a whole-body humanoid loco-manipulation benchmark for household tasks, with the first thorough comparison of learned policies vs LLM-written robot policies.

Key results: GPT-6 Astra lets coding agents write a humanoid controller from just 1 demo, reaching 65% success — surprisingly far, in a regime where learned policies are useless. But π0.5 trained on 60 demos still leads at 75%, especially on precise object rearrangement.

Verdict: learned policies (from scratch or VLAs) remain the more general solution, handling a broader range of tasks at the cost of more demos.

Original post →

More from Embodied

Embodied channel →