Robots get their GPT-3 moment: In-Context Learning lets them learn new tasks from one demo

量子位 · wechat · 2026-09-06

A deep-dive on In-Context Learning spreading to embodied AI: Skild AI's S1 watches one task demo and completes new 10-minute, multi-step tasks without fine-tuning, while Generalist AI's GEN-1.5 does one-shot learning in seconds. Video context reportedly beats language-only prompts by 7x on unseen tasks.

The piece traces the path from GPT-3's ICL to RoboTTT (Fei-Fei Li, Jim Fan, Yuke Zhu) bringing context scaling to robot policies, and interviews Gao Yuxiang, founder of Chinese startup COCO Matrix. He argues that scaling data alone hasn't delivered generalization; the hard parts are hierarchical visual understanding, missing multimodal data (touch/proprioception), and long-horizon memory via KV cache compression. The company bets on "strong understanding, light generation": task-conditioned visual representations that let a 60M action head beat a 1.1B baseline. Next challenges include defining human-teaching interaction datasets.

Original post →

More from Embodied

Embodied channel →