Cornell's Proxy Policy Steering Adapts Frozen Robot Policies Without Touching Weights
philfung · x · 2026-09-28
A Cornell team (including Weichiu Ma and Kuan Fang) proposes Proxy Policy Steering (PPS), an inference-time method that adapts a large frozen robot policy to new tasks from few demonstrations without modifying its weights.
It trains two small networks: a reference proxy distilled from the frozen base policy, and a task proxy fine-tuned from that reference on task demonstrations. At sampling time it steers the base model's diffusion trajectory with a velocity-space residual: vPPS = vbase + γ(vtask − vref).
Notably, the base model's parameters are never modified and aren't even required. It is evaluated on eight real-world manipulation tasks (coffee brewing, flower insertion, jeans folding, shoe retrieval, tissue wiping, etc.) and four simulation tasks.
More from Embodied
- Engineer who 3D-printed carbon fiber for Porsche is now building a flying car — IsaiahBallah · 2026-09-28
- Microbots Built on 20-Year-Old 55nm Process Float Free After Silicon Etch — ctjlewis · 2026-09-28
- Waymo First Ride in Austin: Daily Tesla FSD User Says the Robotaxi Experience Was Mind-Blowing — lydiahallie · 2026-09-28
- Engram is an offline sampler that turns AI hallucinations into music — The Verge AI · 2026-09-28
- Mark Cuban: humanoid robots will fail in 5-10 years, purpose-built forms will win — rohanpaul_ai · 2026-09-28
- Dressing a humanoid robot turns out to be designing a Soft Goods subsystem — yongqianme · 2026-09-28