Hunyuan Research: Batch-Size Scaling with LR Retuning Boosts PPO Throughput 2.29x

TencentHunyuan · x · 2026-09-24

Tencent Hunyuan's new research extends classical critical-batch-size theory to online LLM RL, where the model generates its own training data and rollout generation and training scale differently.

Directly relevant to compute efficiency in large-scale RL training.

Related event: Tencent Hunyuan Extends Critical Batch Size Theory, Boosting LLM RL Throughput(2 posts)→

Original post →

More from Infra

Infra channel →