Adelaide's DCSD Decouples Credit Direction and Magnitude in Self-Distillation, Beats GRPO Across 11 Benchmarks
AdelaideUniversity · hf · 2026-09-30
Researchers from the University of Adelaide propose Decoupled Credit Self-Distillation (DCSD), addressing a fundamental coupling between step-level credit direction and magnitude under teacher supervision in self-distillation.
- The issue: RLVR gives reliable trajectory-level credit while OPSD offers dense token-level supervision, but direction and magnitude become vulnerable to teacher errors and preference variance.
- Method: belief-margin probing determines credit direction; marginal information gain quantifies credit magnitude, enabling step-to-token credit assignment.
- Results: best overall scores vs GRPO, OPSD, RLSD, and RLCSD across 11 benchmarks; +8.45 points on mathematical reasoning and +7.01 on multimodal reasoning over base models, correcting credit direction for 6% of tokens and reducing token credit magnitude 1.5x.
More from Research
- ReScraper: a 0.6B model replaces heuristic data-cleaning stacks, boosting pretraining up to 4.7% — XiongChenyan · 2026-09-30
- Microsoft Research finds LLMs show Dunning-Kruger-style overconfidence in coding — burkov · 2026-09-30
- PixAI launches anime foundation model Tsubaki.3, open-sources Tagger 1.0 with tech report — Level-Ninja-2492 · 2026-09-30
- Lean creator Leonardo de Moura on AI proofs: the Collatz exploit shows verified checkmarks can lie — Machine Learning Street Talk · 2026-09-30
- LLM-42 Paper at SOSP 2026 Brings Deterministic LLM Inference via Verified Speculation — tianyin_xu · 2026-09-30
- Frontier AI Is a Set, Not a Point: Jagged Capabilities May Be the Steady State — vsikka · 2026-09-30