Compute shift to post-training narrows open-closed model gap
A widely discussed essay argues that the compute shift from pretraining to post-training explains the closing gap between open-weight and proprietary models, since post-training advances are far easier to distill and harder to detect.
2026-09-15 ~ 2026-09-15 · 3 related posts
- Post-training shift makes model distillation nearly impossible to detect, researcher argues — maksym_andr · 2026-09-15
- Post-training compute shift makes distillation of frontier models nearly impossible to prevent — maksym_andr · 2026-09-15
1 near-duplicate retellings: maksym_andr