Tencent RL study: larger batches speed up LLM RL if throughput gains beat sample-efficiency loss
CurieuxExplorer · x · 2026-09-25
Tencent's RL team examined how batch size scaling affects LLM reinforcement learning efficiency.
- Larger batches cut time-to-target performance, provided hardware throughput gains outweigh sample-efficiency losses.
- Using a square-root learning rate rule, per-sample learning progress stays stable across a range of batch sizes.
- On fixed hardware, bigger batches are a practical way to lower RL costs and iterate faster.
More from Infra
- Protein design team secures tons of GPUs and lab equipment in single-day buildout — nlarusstone · 2026-09-25
- Speculation: OpenAI's efficient models free compute for bigger ones; Opus 5.5 uses 4x output tokens — haider1 · 2026-09-25
- AI chip startup DensityAI raising hundreds of millions at $10B valuation on AWS deal — steph_palazzolo · 2026-09-25
- xAI said to hit 1M GPUs next week, 1.44M by year-end with 660k GB300s in 90 days — ns123abc · 2026-09-25
- Epoch AI chart shows AI price per intelligence level falling faster than any prior tech — jeff_weinstein · 2026-09-25
- Agents need a global bus: streaming log could be the next agent infra primitive — madhavsinghal_ · 2026-09-25