Emergent Phenomena in Trillion-Parameter Zero RL

inclusionAI · hf · 2026-07-16

This research scales zero RL to 1T parameters, focusing not merely on boosting scores but on observing training dynamics and emergent behaviors at massive scales.

The author notes that naively scaling up training leads to poor readability, repetitive tokens, and non-adaptive reasoning depth. To counter this, they introduced a stable and efficient training pipeline featuring:

The paper presents three core conclusions:

On 7 math benchmarks, Ring-2.5-1T-Zero demonstrates competitive performance. The author also proposes a structured framework to evaluate CoT quality across three dimensions: comprehensibility / reproducibility / efficiency.

Related event: Ring-Zero: Scaling Zero RL to a Trillion Parameters(4 posts)→

Original post →

More from Research

Research channel →