Adobe et al. introduce SpyRL, solving subjective RL rewards via a spy game

burkov · x · 2026-08-16

Reinforcement learning (RL) excels in tasks with deterministic success checks, like math or code, but struggles with summarization or creative writing where there is no single correct answer and training relies on human ratings or AI judges.

A joint work from Adobe, Amazon, Duke University, and others proposes a different route: instead of building a better judge, it changes the training task so that an exact answer exists by construction. The main example, SpyRL, involves multiple model copies performing the same task, with one receiving incomplete information; afterward, the models identify the participant with the missing info.

Since the "spy" identity is fixed beforehand, the training signal can be checked exactly, potentially overcoming the bottleneck of obtaining accurate reward signals in subjective tasks.

Original post →

More from Research

Research channel →