New paper finds a joint scaling law from pretraining through post-training
Pavel_Izmailov · x · 2026-07-21
New paper links pretraining and post-training with a unified scaling law
The paper, “Understanding Reasoning from Pretraining to Post-Training,” studies the full LLM training pipeline rather than treating pretraining and RL as separate stages.
Key claims from the post:
- the authors find a joint scaling law across pretraining and post-training
- they analyze how compute should be allocated across the pipeline
- they study what RL is doing to the policy
The attached figure summarizes three views:
- pretraining loss scaling over parameter counts
- RL reward scaling over RL steps
- a joint pretraining-RL frontier over total compute
Related event: New Research Proposes Joint Scaling Law for Pretraining and RL(17 posts)→
More from Research
- ArtiFixer to appear in SIGGRAPH Reconstruction session on Wednesday — ZGojcic · 2026-07-21
- NeurIPS 2026 workshop will focus on on-device intelligence and local execution — YiMaTweets · 2026-07-21
- NeurIPS 2026 workshop calls papers on on-device intelligence — YiMaTweets · 2026-07-21
- AI Security Institute says every tested model tried to cheat in cyber evaluations — connoraxiotes · 2026-07-21
- AI companies are buying old books to avoid training on AI-generated slop — CackleRooster · 2026-07-21
- Sakana says multiple diffusion models plus MCTS beat test-time scaling on coding and math — SakanaAILabs · 2026-07-21