GPT-Policy teaches robots from human video demos via in-context learning, no retraining
KyeGomezB · x · 2026-09-19
The paper "In-Context Robot Learning with VLM Agents" introduces GPT-Policy, a framework that turns robot learning into prompting:
- Core idea: a frozen commercial VLM (e.g. GPT-6 Astra) adapts to new tasks purely via in-context learning—no gradient updates or task-specific parameter changes.
- Architecture: a context compiler preserving task-relevant visual transitions, a VLM proposing robot-tool actions, and a constrained controller that verifies and executes each action and reports outcomes.
- Key findings: on real robots, human video demonstrations (without robot action labels) improve task completion; aligned action references yield further gains on contact-sensitive tasks.
- Significance: brings few-shot learning into the physical world—teaching a robot a new behavior could look like showing it an example.
Authors span Shanghai Innovation Institute, HUST, FDU and other institutions.
Related event: GPT-Policy: Robots Learn New Tasks via VLM In-Context Learning(2 posts)→
More from Embodied
- Robots are easier to parallelize than cars and can help build themselves — teortaxesTex · 2026-09-19
- Robot fight PR spat: cixliv accused of hiding behind 'the influencer wanted it' — KyleMorgenstein · 2026-09-19
- booster_mjlab: open-source sim-to-real training stack for the Booster K1 humanoid — kevin_zakka · 2026-09-19
- Astra Solves in One Shot What a Robotics Team Spent Two Years Building — YuXiang_IRVL · 2026-09-19
- Sentdex pours cold water on 'rebuild Tesla FSD in an hour' demo: the hard part is perception — teortaxesTex · 2026-09-19
- First human vs robot fight: influencer Frankie Lepenna takes on Terminator robot tonight — cixliv · 2026-09-19