Study: distillation doesn't guarantee better RL outcomes; scale matters
baselabs' study finds that stronger distillation starting points don't guarantee better post-RL accuracy; the benefit of distillation before RL depends on model scale.
2026-09-24 ~ 2026-09-24 · 2 related posts
- Does a Distilled Starting Point Actually Help After RL? A Cross-Size Study — lateinteraction · 2026-09-24
- Study: Higher pre-RL accuracy from distillation doesn't guarantee higher post-RL accuracy — teortaxesTex · 2026-09-24