Absolute Zero: Self-Play RL Reasoning With Zero Human Data

cephaloform · x · 2026-09-04

The arXiv paper "Absolute Zero" proposes a new RLVR paradigm where a single model proposes tasks that maximize its own learning progress and improves reasoning by solving them — no external data needed. The Absolute Zero Reasoner (AZR) uses a code executor to both validate proposed code-reasoning tasks and verify answers as a unified verifiable reward, self-evolving its curriculum. Fans describe it as "a tiny code-capabilities arms race you can watch run for half an hour."

Original post →

More from Research

Research channel →