Why the Open-Weight Gap Is Shrinking: Post-Training Compute Is Easier to Distill
maksym_andr · x · 2026-09-15
The author offers a mechanistic explanation for why the gap between open-weight and publicly available proprietary models has narrowed: the compute shift from pretraining to post-training means improvements now flow through channels that are easier to distill. A very capable model in a domain can be used to create RL environments for that domain, and this kind of distillation is hard to detect since environment generation is stealthier than copying raw outputs.
Related event: Compute shift to post-training narrows open-closed model gap(3 posts)→
More from Models
- Grok 4.7 Misses Target Again; 2.5T-Parameter Grok 4.8 Finishes Training — eyishazyer · 2026-09-15
- Writers Report Gemini Flash 3.8 Suffers Severe Long-Context Rot Despite 1M Token Window — Quenty1 · 2026-09-15
- Ai2 Chair's 3 Predictions: Open-Source AI Goes from Ideology to Infrastructure in 18 Months — billhilf · 2026-09-15
- Surgical VLM Leaderboard: All Frontier Models Fall Far Short of Specialized Models — ddonoho · 2026-09-15
- Why Surgery Benchmarks Reward Fine-tuned Small Models While Math Benchmarks Don't — ddonoho · 2026-09-15
- Claude Opus 5.2 spotted in grayscale testing on Claude Code, seemingly skipping 5.1 — Angaisb_ · 2026-09-15