Post-training shift makes model distillation nearly impossible to detect, researcher argues
maksym_andr · x · 2026-09-15
The author argues the compute shift from pretraining to post-training explains why open-weight models keep closing the gap: post-training gains are far easier to distill.
Key points:
- Access to a capable model in a domain lets you generate RL environments for that domain, effectively distilling it;
- Such distillation is hard to detect because generating environments looks like legitimate use;
- Only a large volume of environment-generation requests might give it away, but detection is post-hoc and multi-account evasion is trivial;
- Conclusion: preventing distillation in full generality seems nearly impossible.
Related event: Compute shift to post-training narrows open-closed model gap(3 posts)→
More from Research
- FlyWire publishes DIY guide to simulate a fly brain with ~160k neurons — patrickmineault · 2026-09-15
- Stanford and MIT paper: the code harness around an LLM can swing benchmark results up to 6x — burkov · 2026-09-15
- Close to a huge math breakthrough, then scooped by AI: what it means for open science — ScottNover · 2026-09-15
- ECCV 2026 paper ART fixes complex makeup transfer, releases first 2K dataset — jiqizhixin · 2026-09-15
- Why mathematicians resist AI proofs — and why it's not just gatekeeping — rbhar90 · 2026-09-15
- CCN2026 GAC debate recording: is the NeuroAI approach inevitable for understanding the brain? — aran_nayebi · 2026-09-15