Pretraining efficiency gains come mostly from data, but end-to-end task gains spread across RL and systems
eigenron · x · 2026-10-10
eigenron cites Dwarkesh's recent experiments to argue that most compute-efficiency gains in pretraining have come from better data rather than model architecture improvements. However, for actual end-to-end task performance, the gains are much more evenly distributed between RL, systems/engineering, and data, rather than being overwhelmingly data-driven.
More from Research
- Regents Labs launches Paper Pro Daily, an AI paper-a-day digest built with ChatGPT Pro — seanwbren · 2026-10-10
- KKT Points in Imperfect-Recall Games Correspond to CDT+GT Optimal Policies — jessi_cata · 2026-10-10
- Comprehensive sweep finds the best optimizer changes with batch size — aaron_defazio · 2026-10-10
- Dino Forcing paper predicts DINO reps instead of co-denoising, halving training epochs on ImageNet — kastnerkyle · 2026-10-10
- Community + AI agents push integer multiplication bounds to κ≈6.83e-4, hitting a structural wall — ChrSzegedy · 2026-10-10
- Microsoft's CABRA shows coding agents bottleneck on code understanding, not edit size — rohanpaul_ai · 2026-10-10