RLVR Can't Prove What Its Verifier Can't: The ZFC Ceiling Argument

Liu_eroteme · x · 2026-09-29

A debate on the ceiling of RLVR (RL with verifiable rewards): the poster argues both pretraining and posttraining are limited — posttraining can only approach the best policies an external verifier can verify, so models will end up "equal to or slightly better than the best humans at any task" rather than superhuman.

In follow-up he sharpens the point: if the verifier is itself a subset of ZFC set theory, RLVR has no reward signal for theorems independent of ZFC — no possible policy gradient can distinguish them from "impossible." The core claim: RLVR is capped by the expressive power of its verifiers, so superhuman mathematical reasoning can't come from RLVR alone.

Related event: RLVR Is Capped by Verifiers, Limiting Path to Superhuman AI(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →