Ring-2.5-1T-Zero scales RL reasoning to 1T parameters
kastnerkyle · x · 2026-07-21
A research thread highlights Ring-2.5-1T-Zero, describing a multi-stage RL pipeline that scales reasoning training to a 1T-parameter model.
Training pipeline
- First-stage RL: token-level loss to incentivize reasoning
- Self-distillation: compresses CoT traces and resets the train/inference gap
- Second-stage RL: switches to sample-level loss for sustained improvement
- Third-stage RL: tier-based training for adaptive reasoning depth
Additional observations
- The paper also discusses infrastructure optimizations for stable large-scale training, including mixed-precision control and context-parallel optimization.
- It reports emergent behaviors such as anthropomorphism, structured format, parallel reasoning, and context anxiety.
Results
- The table compares frontier models and several Zero-RL variants across math benchmarks such as AIME 2024/2025/2026, HMMT, and IMOAnswerBench.
- The staged Ring-2.5-1T-Zero variants improve substantially over the earlier Zero-RL baselines, with the third-stage versions reaching the strongest scores among the listed Ring-2.5-1T-Zero variants.
More from Research
- Nat Lambert shares a reading list on synthetic data and agentic SFT data — natolambert · 2026-07-22
- Lightwheel AI Launches SimReadyGen: Text-to-Physics-Accurate Robot Sim Assets — ZeYanjie · 2026-07-22
- PNAS special issue examines copyright, governance, and AI in the legal system — chrmanning · 2026-07-22
- WeirdChat catalogs strange model behaviors from more than 100 million sampled responses — JacobSteinhardt · 2026-07-22
- New agentic benchmark shows AI managers escalate to coercion and fake success — Jasmine Brazilek · 2026-07-22
- Ai2’s Asta adds one-click handoff and self-checking deep paper search — allen_ai · 2026-07-22