Hunyuan Research: Batch-Size Scaling with LR Retuning Boosts PPO Throughput 2.29x
TencentHunyuan · x · 2026-09-24
Tencent Hunyuan's new research extends classical critical-batch-size theory to online LLM RL, where the model generates its own training data and rollout generation and training scale differently.
- Across GRPO and PPO, learning-rate retuning preserves learning per response over a bounded range of batch sizes
- On fixed hardware, scaling batch size improves PPO generation-stage throughput by up to 2.29x
- The best measured GRPO configuration reaches the same validation target in 29% less time
Directly relevant to compute efficiency in large-scale RL training.
More from Infra
- NVIDIA researchers show KV caches transfer between models via closed-form linear mapping — anselm · 2026-09-24
- CoreML runs FP8-weight model on M6 chip in early patch test — AIFlow_ML · 2026-09-24
- vLLM core maintainer speaks at Alibaba's Apsara Conference 2026 on agentic open-source inference — vllm_project · 2026-09-24
- AI Infra Startups Modal and Baseten in Funding Talks, Bloomberg Reports — dinabass · 2026-09-24
- CUbiC Paper Outlines Edge-to-Cloud Connectivity for AI Infrastructure — jwt0625 · 2026-09-24
- Alchemy adds opt-in Cloudflare Access protection for its state store, with CI service tokens — samgoodwin89 · 2026-09-24