From Gradients to Capabilities: How Adam and BF16 Flatten Multi-Teacher Distillation Signals
UIUC-CS · hf · 2026-10-06
Studying Qwen3-1.7B with four same-init RL-trained domain teachers (plus SmolLM3-3B diagnostics), the paper dissects how teacher signals shape parameter updates in multi-teacher on-policy distillation:
- Loss averaging implicitly weights responses: token averaging favors longer responses, and equalizing domain contributions preserves this within domains
- Adam's first moment flattens update differences: cosine similarity 0.83 between teachers, 0.96 between averaging rules, despite raw gradient differences
- BF16 rounding hides small changes: 97% of FP32 master weights differ from init, but only 7–11% of BF16 weights do
- The top-64 intersection KL gradient closely matches the full-vocabulary gradient, but task impact hinges on averaging: math accuracy is 2.6 points higher than sampled-token PG under response averaging, 2.1 points lower under global token averaging
More from Research
- Subsampling and extrapolation keep the Mandelbrot area estimate unbiased near the boundary — geoffreyirving · 2026-10-06
- New estimate sits 6.5e-9 below Hsing Lo's 2025 value; reproduction suggests the gap is a fluctuation — geoffreyirving · 2026-10-06
- Claude-assisted CUDA compute pins Mandelbrot set area to 1.506591883653, 60x tighter than 2012 record — geoffreyirving · 2026-10-06
- Group-Evolving Agents: a new paradigm where the unit of agent self-improvement is a group — xwang_lk · 2026-10-06
- SLIM paper at COLM: design principles for long-horizon agentic search systems — xiye_nlp · 2026-10-06
- The Nobel optogenetics drama: forgotten inventor Zhuo-Hua Pan had the stronger claim — _onionesque · 2026-10-06