Post-training compute shift makes distillation of frontier models nearly impossible to prevent
maksym_andr · x · 2026-09-15
The shift of compute from pretraining to posttraining explains why the gap between open-weight and proprietary models has shrunk: post-training improvements are much easier to distill.
Key argument:
- Access to a very capable model in some domain can be used to create RL environments in that domain, enabling distillation
- This is hard to detect from the outside: generating an environment looks like legitimate use
- Only a large volume of environment-related requests might give it away — and only post-hoc, after too many environments already exist
- Dedicated organizations can trivially create many accounts to evade detection
Conclusion: preventing distillation in full generality seems nearly impossible.
Related event: Compute shift to post-training narrows open-closed model gap(3 posts)→
More from Models
- Building your own agent harness beats bloated off-the-shelf frameworks on cost, argues researcher — omarsar0 · 2026-09-15
- Agent Arena: DeepSeek V4.1 Flash hits Pareto frontier at $0.06/task with +4.87% net improvement — arena · 2026-09-15
- 23 Days Without Claude Code: Dev Says Codex Works Better With OSS, Kimi K3 Unbeaten at Coding — Yuchenj_UW · 2026-09-15
- Marigold-V2 depth estimation demo trends on Hugging Face Spaces — toshas · 2026-09-15
- OpenAI has hundreds of contractors reading and rating your ChatGPT chats — The Decoder · 2026-09-15
- Hugging Face Rounds Up Which Open LLMs Are Best for On-Device Inference — NielsRogge · 2026-09-15