MIT Research: Fixing RLVR Diversity Collapse with Adversarial Discriminators

dair_ai · x · 2026-07-04

DAIR.AI recommended an MIT study on Reinforcement Learning with Verifiable Rewards (RLVR). Because RLVR only optimizes objectively scorable dimensions, it leads to a quiet collapse in style, structure, and diversity, while encouraging reward hacking. The work introduces an adversarial discriminator trained on human demonstrations to act as a proxy for human output distribution. This forces the generator to optimize for both task accuracy and "humanness," proving effective in tasks like bug fixing and story generation.

Original post →

More from Research

Research channel →