Comparing Peer Distillation vs Post-Training

maxsloef · x · 2026-07-18

The author notes that fable's literature review mostly uncovered papers on "strong teacher to weak student" distillation. What they actually want to explore is a different scenario: instead of traditional teacher-student distillation, distilling the rollouts of a similarly performing peer model into another model, to see if it matches the performance of "standard post-train first, then distill". This concept is compared to setups like k3/opus.

Related event: Exploring Peer Distillation and Post-Training Between Equal Models(2 posts)→

Original post →

More from Research

Research channel →