Pedagogical RL: privileged info should actively sample rollouts, not just score them

lateinteraction · x · 2026-09-07

A technical thread on Pedagogical RL. The quoted post argues typical RL algorithms and on-policy distillation are blind samplers: privileged info scores rollouts but isn't used to find them. The reply adds that SDPO's failure was predictable — it only rescores rollouts with privileged info instead of using it to guide sampling — so when that info is high-entropy/unexpected, stumbling upon successful rollouts is hopeless. The proposal: use privileged info to actively sample the rollouts RL wishes it could stumble upon.

Original post →

More from Research

Research channel →