Intention Distillation lifts GR00T success from 64% to 85% on SimplerEnv
POSTECH · hf · 2026-08-31
POSTECH proposes Intention Distillation (INDI): during training a frozen VLM teacher distills behavior-level intent from the instruction, a coarse action summary, and execution video; the deployed VLA recovers this multimodal intent at an intermediate decoder layer to organize action prediction.
Results:
- GR00T-N1.7: 64.3% → 84.7% on SimplerEnv-Bridge
- 64.1% → 70.3% on RoboCasa Kitchen
- Consistent gains for π₀.₅ on both benchmarks
- Real-world average success 62.0% → 68.7%, up to 12pp on longer-horizon tasks
Analyses show the recovered latent captures behavior objective and execution progress.
More from Embodied
- Open Source Project: Distill openpilot to Train Custom Autonomous Driving Models — yassineyousfi_ · 2026-09-01
- China scales cheap humanoid bodies first, letting intelligence catch up later — VraserX · 2026-09-01
- NUS MAGIC Lab recruiting: Focusing on Embodied AI research — DJiafei · 2026-09-01
- Why Microduck is Winning: Cuteness Overload and Sim2real Tech — RachelVT42 · 2026-09-01
- Seeking recs: What are the most underrated AI-native hardware gadgets? — thisiskp_ · 2026-09-01
- ICE to use Boston Dynamics robot dogs in operations — r_singh · 2026-09-01