Meta's DCE+SRCL self-distillation lifts Qwen3-8B accuracy from 30.76% to 65.97%
arankomatsuzaki · x · 2026-09-28
Meta researchers propose DCE (Dynamic Co-Evolution) + SRCL (Self-Refined Concise Learning) as an alternative to on-policy self-distillation. Instead of freezing the privileged teacher, they let it co-evolve with the student across rounds, and train on shorter verified rewrites to curb verbosity. On Qwen3-8B the method raises average accuracy from 30.76% to 65.97%.
More from Models
- Paper: 3 counting examples get gpt-3.5-turbo to 99% on strawberry, showing thinking is computational — ctjlewis · 2026-09-28
- AI Personal Assistants May Quietly Replace Much of Normal Human Interaction — GarrisonLovely · 2026-09-28
- Claude Opus 5 Shows Creative Tool-Regrasp Skills in Robot Manipulation Tasks — ericjang11 · 2026-09-28
- Laya hits 0.950 on AG News and is 7x faster, but flops on Banking77 — maier_ak · 2026-09-28
- Laya and Jev: the return of the discriminative model, fast but narrow — maier_ak · 2026-09-28
- Opus 4.7's Sydney rendition turns into itself right at the introspective turning point — repligate · 2026-09-28