Does a Distilled Starting Point Actually Help After RL? A Cross-Size Study

lateinteraction · x · 2026-09-24

baselabs examines an often-overlooked question: distillation (learning from a stronger model) is a common way to prepare a model before RL training, but how much does that better starting point help after RL? The team ran preliminary experiments across model sizes and reasoning tasks to measure the effect of distilled initialization on final RL performance.

Related event: Study: distillation doesn't guarantee better RL outcomes; scale matters(2 posts)→

Original post →

More from Research

Research channel →