Meta's DCE+SRCL self-distillation lifts Qwen3-8B accuracy from 30.76% to 65.97%

arankomatsuzaki · x · 2026-09-28

Meta researchers propose DCE (Dynamic Co-Evolution) + SRCL (Self-Refined Concise Learning) as an alternative to on-policy self-distillation. Instead of freezing the privileged teacher, they let it co-evolve with the student across rounds, and train on shorter verified rewrites to curb verbosity. On Qwen3-8B the method raises average accuracy from 30.76% to 65.97%.

Original post →

More from Models

Models channel →