MaD-RL author explains why reward-only RL can't target output mixture ratios

nagpalchirag · x · 2026-09-24

In a follow-up on the paper thread, the author clarifies that GRPO-style RL is known to concentrate probability on particular modes, and reward-only maximization gives no prompt-specific distributional control. In fact, correctness-only RL can actively hurt response diversity. Entropy bonuses restore spread but cannot specify an arbitrary target mixture — the gap MaD-RL aims to fill.

Related event: Meta Team Proposes MaD-RL to Calibrate LLM Output Distributions with RL(2 posts)→

Original post →

More from Research

Research channel →