Emergent Phenomena in Trillion-Parameter Zero RL
inclusionAI · hf · 2026-07-16
This research scales zero RL to 1T parameters, focusing not merely on boosting scores but on observing training dynamics and emergent behaviors at massive scales.
The author notes that naively scaling up training leads to poor readability, repetitive tokens, and non-adaptive reasoning depth. To counter this, they introduced a stable and efficient training pipeline featuring:
- clipped importance sampling
- training-inference ratio correction
- mixed-precision control
The paper presents three core conclusions:
- Scaling to 1T parameters significantly improves sample efficiency and performance ceilings
- The training process transitions from a "discovery phase" to a "refinement phase"
- The model spontaneously exhibits higher-level cognitive behaviors, such as anthropomorphism, structured outputs, self-verification, parallel reasoning, and context anxiety
On 7 math benchmarks, Ring-2.5-1T-Zero demonstrates competitive performance. The author also proposes a structured framework to evaluate CoT quality across three dimensions: comprehensibility / reproducibility / efficiency.
Related event: Ring-Zero: Scaling Zero RL to a Trillion Parameters(4 posts)→
More from Research
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11
- Harry Collins: LLMs can't do frontier science because they can't invent new language — whoamisri · 2026-09-11
- The Waymo effect: how AI is quietly making research less collaborative — JohnHammersley · 2026-09-11
- Causal-only attention for non-generative tasks is wasteful, argues HF engineer — antoine_chaffin · 2026-09-11
- Catholic University of Chile researcher: scaling AI feedback is key to sustainable medical education — julianvarascom · 2026-09-11
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11