AVSD accepted at NeurIPS 2026: adaptive multi-view self-distillation improves LLM reasoning RL

mohitban47 · x · 2026-09-30

AVSD (Adaptive-View Self-Distillation) has been accepted to NeurIPS 2026. It tackles better on-policy token-level learning signals for LLM reasoning RL, where sparse binary rewards are the bottleneck. Existing on-policy self-distillation methods condition the teacher on a single privileged view (full solution, partial rationale, answer-only, reference code, etc.), but no single view is consistently best, and views can introduce teacher-specific artifacts from information unavailable to the student.

AVSD's key idea: useful supervision comes from both what teachers agree on and information unique to individual teachers. Multiple views of privileged information (hints, partial/full solutions, execution outputs) each induce a different teacher distribution, and the method adaptively combines them to produce better token-level signals for math and code reasoning.

Original post →

More from Research

Research channel →