SFT+RL pipeline choice makes the lazy-distillation theory unlikely, argues dev

JoshPurtell · x · 2026-09-03

Follow-up in JoshPurtell's distillation debate thread:

Complements the earlier post: the model's gains likely come from RL, not distillation.

Related event: JoshPurtell breaks down the model distillation debate: task type determines risk and feasibility(8 posts)→

Original post →

More from Models

Models channel →