RLSVR: Task Transformation Enables Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

burny_tech · x · 2026-08-15

The paper introduces RLSVR, extending RLVR beyond math and code by making rewards verifiable through task transformation. Using SpyRL, it employs information-asymmetric self-play: one agent gets degraded context, all agents perform the same task, then vote on the spy. The known spy identity gives an exact detector reward, while vote counts become a relative performer reward. This converts subjective output quality into a self-generated RL signal without reward models or external judges.

Original post →

More from Research

Research channel →