Post-training compute shift makes distillation of frontier models nearly impossible to prevent

maksym_andr · x · 2026-09-15

The shift of compute from pretraining to posttraining explains why the gap between open-weight and proprietary models has shrunk: post-training improvements are much easier to distill.

Key argument:

Conclusion: preventing distillation in full generality seems nearly impossible.

Related event: Compute shift to post-training narrows open-closed model gap(3 posts)→

Original post →

More from Models

Models channel →