RLVR Is Guaranteed to Sigmoid: Verifiers Bounded by Human-Written Math Can't Escape ZFC

Liu_eroteme · x · 2026-09-29

A substantive technical take: current RLVR relies on human-written verifiers (and agent-written ones inherit the same limits), so the approach is guaranteed to follow a sigmoid curve. If the verifier is a subset of ZFC, RLVR cannot train an agent to prove theorems beyond ZFC — no policy gradient exists to distinguish 'impossible' from 'hard'. The author sees RLIR on real-world interaction as the only unbounded path, but it's ratelimited by real time.

Related event: RLVR Is Capped by Verifiers, Limiting Path to Superhuman AI(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →