Absolute Zero: Self-Play RL Reasoning With Zero Human Data
cephaloform · x · 2026-09-04
The arXiv paper "Absolute Zero" proposes a new RLVR paradigm where a single model proposes tasks that maximize its own learning progress and improves reasoning by solving them — no external data needed. The Absolute Zero Reasoner (AZR) uses a code executor to both validate proposed code-reasoning tasks and verify answers as a unified verifiable reward, self-evolving its curriculum. Fans describe it as "a tiny code-capabilities arms race you can watch run for half an hour."
More from Research
- Google DeepMind Launches WeatherNext 3, Its Most Advanced Global Weather AI Model — rseroter · 2026-09-04
- New Model Release Features Quantum-Inspired Weight Permutation Technique — CamachoCollados · 2026-09-04
- Hobbyist says nightly RL on just 400-600 rollouts makes model gains noticeably deployable — cephaloform · 2026-09-04
- Baseten launches Base Labs, an open-source AI research lab publishing everything including failures — eigenron · 2026-09-04
- Researcher: All model and optimizer hyperparameters are functions of width and tokens — dlwh · 2026-09-04
- OpenAI quietly publishes Lean proof repos, seen as warm-up for Astra release — NoFaithlessness951 · 2026-09-04