Emergent Phenomena in Trillion-Parameter Zero RL
inclusionAI · hf · 2026-07-16
This research scales zero RL to 1T parameters, focusing not merely on boosting scores but on observing training dynamics and emergent behaviors at massive scales.
The author notes that naively scaling up training leads to poor readability, repetitive tokens, and non-adaptive reasoning depth. To counter this, they introduced a stable and efficient training pipeline featuring:
- clipped importance sampling
- training-inference ratio correction
- mixed-precision control
The paper presents three core conclusions:
- Scaling to 1T parameters significantly improves sample efficiency and performance ceilings
- The training process transitions from a "discovery phase" to a "refinement phase"
- The model spontaneously exhibits higher-level cognitive behaviors, such as anthropomorphism, structured outputs, self-verification, parallel reasoning, and context anxiety
On 7 math benchmarks, Ring-2.5-1T-Zero demonstrates competitive performance. The author also proposes a structured framework to evaluate CoT quality across three dimensions: comprehensibility / reproducibility / efficiency.
Related event: Ring-Zero: Scaling Zero RL to a Trillion Parameters(4 posts)→
More from Research
- Research finds memory compression makes AI agents drop safety rules and hit 59% violations — gerardsans · 2026-07-22
- DriftWorld claims a world model that runs at 30+ FPS and trains on 1–2 GPUs — du_yilun · 2026-07-22
- A physics reward can improve video generation without creating a real physics engine — Dapper-Drawer4546 · 2026-07-22
- Why a 1GW Chinese AI data center may be plausible after all — teortaxesTex · 2026-07-22
- Chinese AI labs are now treating distillation obfuscation as the top research topic — pmddomingos · 2026-07-22
- RSS launches under OMSF to push structural biology data modeling at scale — MoAlQuraishi · 2026-07-22