DiffGate applies teacher guidance only to failed rollouts, boosting GRPO pass@8 by up to 5.7 points
Snapchat · hf · 2026-10-07
A new Hugging Face paper introduces DiffGate, a difficulty-gated objective for on-policy distillation that combines GRPO with selective teacher guidance.
Core idea
- On-policy distillation (OPD) gives dense local supervision but is weakly aligned with rollout correctness; RLVR/GRPO gives outcome-level supervision but suffers sparse rewards and coarse credit assignment.
- DiffGate applies teacher supervision only to failed trajectories, scaled by group difficulty and smoothly bounded; the verifier decides which trajectories get guidance, while the teacher supplies dense token-level update directions within them.
Results
- On Qwen3-0.6B / 1.7B students, code avg@8 improves over GRPO by +1.7 / +1.8 points, and pass@8 by +1.6 / +5.7 points.
- On math, avg@8 stays within 0.5 points of GRPO while pass@8 improves by +1.1 / +3.9 points.
- pass@8 improves across all four model–domain settings, indicating better solution coverage.
More from Research
- Fixed token codes suffice: 1.7B LM trains without a trainable input embedding table — A. Bochkov · 2026-10-07
- EmbeddingGemma 2 hands-on: 740M multimodal embeddings for search and RAG, runnable on a free T4 — Prompt Engineering · 2026-10-07
- Isomorphic, DeepMind and Meta join DOE-NIH partnership to build an AI model of the cell — snikolov · 2026-10-07
- Bi-manual mobile UMI demo unlocked for robot manipulation data collection — neurosp1ke · 2026-10-07
- Researcher presents Meta-Harness and Combee at COLM 2026, seeks industry roles — Kangwook_Lee · 2026-10-07
- Lampinen: great cultural insights rarely come from a single brain, unlike LLM analogy — AndrewLampinen · 2026-10-07