Joint pretraining-RL scaling says RL share should rise with compute
teortaxesTex · x · 2026-07-20
A scaling-law discussion on joint pretraining and RL suggests that the optimal RL share rises with total compute, but the pretraining allocation itself does not deviate much from Chinchilla-style scaling. The author highlights a synthetic chess setting where reward, pretraining loss, and RL compute are linked, and argues that for pass@1, pretraining affects results in two ways: it lowers downstream loss and improves how efficiently RL compute turns into reward. The post also warns that naively extrapolating the fitted law to very large compute can produce impossible rewards and wildly inflated RL shares, so the shared scaling trend should be kept, but capped conservatively.
Related event: New Research Proposes Joint Scaling Law for Pretraining and RL(17 posts)→
More from Research
- Kimi K3 and Fable 5 now look much closer than the old open-vs-closed gap — FinanceYF5 · 2026-07-21
- uv-scripts/ocr returns to the top of Hugging Face datasets with a JSON model picker — vanstriendaniel · 2026-07-21
- DeepSearch-World trains web agents with 420K verifiable QA tasks — HKUST · 2026-07-21
- GigaAM Multilingual targets low-resource Central Asian ASR with 2M hours of audio — ai-sage · 2026-07-21
- WorldCupArena benchmarks language models on 104 football matches — Zhaokai Wang · 2026-07-21
- Reddit asks whether LLMs need a benchmark for treasure-hunt style reasoning — StrangeOops · 2026-07-21