Freeform Preference Learning Boosts Robot Manipulation Over Baselines by 38 Percentage Points
StanfordAILab · x · 2026-09-10
Chelsea Finn and colleagues presented Freeform Preference Learning for Robotic Manipulation (arXiv:2606.32027) on the RoboPapers podcast.
- Problem: Reward design is a bottleneck in robot RL — sparse rewards give too little signal, and binary preference learning collapses multiple quality notions into one ambiguous label.
- Method: FPL lets annotators define natural-language preference axes (speed, safety, placement quality, carefulness) and give pairwise preferences per axis; annotations train a language-conditioned reward model, which trains a reward-conditioned policy.
- Results: Across 4 real-world and 2 simulated long-horizon manipulation tasks, FPL beats sparse-reward and binary-preference methods by 38 percentage points; it learns dense progress signals without subtask segmentation, shows emergent compositionality, and lets users steer behavior at test time without retraining.
Related event: Freeform Preference Learning Defines Robot Rewards in Natural Language(2 posts)→
More from Embodied
- Madrid to become first European city with autonomous taxi service, CA-plated cars spotted — eherrerosj · 2026-09-10
- Dexory's autonomous robot audits entire warehouses daily, replacing manual stock counts — lukas_m_ziegler · 2026-09-10
- DexGPT: a GPT-written real2sim pipeline turns one monocular hand GIF into 44-DOF dexterous sim, open-sourced — animesh_garg · 2026-09-10
- "A billion robots and a bad firmware upgrade" sparks open-source-only trust debate — tomchapin · 2026-09-10
- Gurman: foldable iPhone to be named iPhone Duo, starting at $2,000 — mitchdeg · 2026-09-10
- BHF Robotics' Gen3 kills weeds by electrocuting roots, no herbicide needed — chrisgrayson · 2026-09-10