RLVR Is Guaranteed to Sigmoid: Verifiers Bounded by Human-Written Math Can't Escape ZFC
Liu_eroteme · x · 2026-09-29
A substantive technical take: current RLVR relies on human-written verifiers (and agent-written ones inherit the same limits), so the approach is guaranteed to follow a sigmoid curve. If the verifier is a subset of ZFC, RLVR cannot train an agent to prove theorems beyond ZFC — no policy gradient exists to distinguish 'impossible' from 'hard'. The author sees RLIR on real-world interaction as the only unbounded path, but it's ratelimited by real time.
Related event: RLVR Is Capped by Verifiers, Limiting Path to Superhuman AI(3 posts)→
More from AGI Musings
- MIT's Daniela Rus: intelligence will live in your pocket, not a data center — MIT_CSAIL · 2026-09-30
- Why evolving systems can never fully self-operate without external supervision — attentionmech · 2026-09-29
- MIT's Cristiano Catalini: never let intelligence become too big to fail — WhatTheLJW · 2026-09-29
- Lenny's Summit takeaway: agent work is lonely, curated IRL events have huge PMF — lennysan · 2026-09-29
- Microsoft Research unveils Project Quine, an AI research system combining biology world model with wet lab — erichorvitz · 2026-09-29
- Autonomous AI agents are killing the ad-supported web — and advertisers are paying for bot views — moderndaydjangoo · 2026-09-29