Post-training shift makes model distillation nearly impossible to detect, researcher argues

maksym_andr · x · 2026-09-15

The author argues the compute shift from pretraining to post-training explains why open-weight models keep closing the gap: post-training gains are far easier to distill.

Key points:

Related event: Compute shift to post-training narrows open-closed model gap(3 posts)→

Original post →

More from Research

Research channel →