Tencent Hunyuan extends critical-batch-size theory to LLM RL: 29% faster GRPO, 2.29× PPO throughput
TencentHunyuan · x · 2026-09-24
Tencent Hunyuan's new research extends classical critical-batch-size theory to online LLM RL, where the model generates its own training data and rollout scaling differs from training scaling. Key findings: across GRPO and PPO, learning-rate retuning preserves learning per response over a bounded batch-size range; on fixed hardware, larger batches boost PPO generation-stage throughput by up to 2.29×; and the best GRPO configuration hits the same validation target in 29% less time.
More from Infra
- Modal explains how it serves trillions of tokens for coding agents — charles_irl · 2026-09-24
- AI Infra Startups Modal and Baseten in Funding Talks, Bloomberg Reports — dinabass · 2026-09-24
- CUbiC Paper Outlines Edge-to-Cloud Connectivity for AI Infrastructure — jwt0625 · 2026-09-24
- Hunyuan Research: Batch-Size Scaling with LR Retuning Boosts PPO Throughput 2.29x — TencentHunyuan · 2026-09-24
- Alchemy adds opt-in Cloudflare Access protection for its state store, with CI service tokens — samgoodwin89 · 2026-09-24
- A Ready-to-Use Prompt That Makes Your Agent Audit Its Own API Bills — gethackteam · 2026-09-24