Tencent Hunyuan extends critical-batch-size theory to LLM RL: 29% faster GRPO, 2.29× PPO throughput

TencentHunyuan · x · 2026-09-24

Tencent Hunyuan's new research extends classical critical-batch-size theory to online LLM RL, where the model generates its own training data and rollout scaling differs from training scaling. Key findings: across GRPO and PPO, learning-rate retuning preserves learning per response over a bounded batch-size range; on fixed hardware, larger batches boost PPO generation-stage throughput by up to 2.29×; and the best GRPO configuration hits the same validation target in 29% less time.

Related event: Tencent Hunyuan Extends Critical Batch Size Theory, Boosting LLM RL Throughput(2 posts)→

Original post →

More from Infra

Infra channel →