PARTS: subtask RL fine-tuning lifts robot success from 32% to 61% with minimal supervision
Sichang Su · hf · 2026-09-22
- Problem: pretrained robot policies often fail at a few critical subtasks of long-horizon jobs; extra full-task demos for SFT are costly, and sparse-reward RL struggles on long horizons.
- Method: PARTS concentrates RL practice on bottleneck subtasks — a frozen pretrained policy supplies nominal actions while agent-generated selectors and success verifiers activate residual corrections with local outcome rewards; online RL plus success-reweighted retraining iterates with redeployed residual policies. Humans only flag bottlenecks at setup and do physical resets.
- Results: full-task success improves from 32%→61% (bimanual YAM) and 50%→95% (Franka), beating prior real-world RL fine-tuning by >25 points under the same rollout budget with far less human involvement.
More from Embodied
- Meta Ray-Ban glasses talking to Muse could become the best consumer AI app overnight — ChrisUniverse · 2026-09-22
- JHU to host Scalable Tactile Sensing for Dexterous Manipulation workshop at IROS 2026 — _krishna_murthy · 2026-09-22
- SLIM-init: line-feature VIO initialization for degenerate motions, accepted to IROS 2026 — zhenjun_zhao · 2026-09-22
- LoG-VGGT uses cross-window attention for memory-efficient long-sequence 3D reconstruction — zhenjun_zhao · 2026-09-22
- VoxelTTO: voxel-aligned feed-forward 3DGS with test-time optimization, 80 GPU hours — zhenjun_zhao · 2026-09-22
- Info3R: information-adaptive test-time training cuts KITTI pose error 1.68x vs LongStream — zhenjun_zhao · 2026-09-22