Tencent Hunyuan Extends Critical Batch Size Theory, Boosting LLM RL Throughput
Tencent Hunyuan extended critical batch size theory to online LLM RL, achieving 29% faster GRPO training and up to 2.29x PPO generation throughput by scaling batch size with learning-rate adjustments.
2026-09-24 ~ 2026-09-24 · 2 related posts
- Tencent Hunyuan extends critical-batch-size theory to LLM RL: 29% faster GRPO, 2.29× PPO throughput — TencentHunyuan · 2026-09-24
- Hunyuan Research: Batch-Size Scaling with LR Retuning Boosts PPO Throughput 2.29x — TencentHunyuan · 2026-09-24