SpyRL: Using 'Who Is the Spy' Mechanics for LLM Self-Improvement on Open-Ended Tasks

Qinsi Wang · hf · 2026-08-03

To break the bottleneck of Reinforcement Learning with Verifiable Rewards (RLVR) outside math and coding, researchers propose RLSVR (Reinforcement Learning with Self-Verifiable Rewards). This paradigm transforms open-ended tasks into verifiable proxy environments that automatically generate reward signals.

The team instantiates this with SpyRL, a multi-agent self-play environment inspired by 'Who Is the Spy?'. Agents receive asymmetric information, complete target tasks, and vote to identify a predetermined spy. Because the spy identity is known, voting outcomes provide fully verifiable rewards, while successful identification correlates with output quality.

Experiments on text summarization, creative writing, and math reasoning show SpyRL outperforms existing self-improvement methods on non-verifiable tasks while maintaining gains on verifiable reasoning tasks.

Related event: New RLSVR Paradigm Enables Open-Ended LLM Self-Correction(4 posts)→

Original post →

More from Research

Research channel →