Why the Open-Weight Gap Is Shrinking: Post-Training Compute Is Easier to Distill

maksym_andr · x · 2026-09-15

The author offers a mechanistic explanation for why the gap between open-weight and publicly available proprietary models has narrowed: the compute shift from pretraining to post-training means improvements now flow through channels that are easier to distill. A very capable model in a domain can be used to create RL environments for that domain, and this kind of distillation is hard to detect since environment generation is stealthier than copying raw outputs.

Related event: Compute shift to post-training narrows open-closed model gap(3 posts)→

Original post →

More from Models

Models channel →