MaD-RL author explains why reward-only RL can't target output mixture ratios
nagpalchirag · x · 2026-09-24
In a follow-up on the paper thread, the author clarifies that GRPO-style RL is known to concentrate probability on particular modes, and reward-only maximization gives no prompt-specific distributional control. In fact, correctness-only RL can actively hurt response diversity. Entropy bonuses restore spread but cannot specify an arbitrary target mixture — the gap MaD-RL aims to fill.
Related event: Meta Team Proposes MaD-RL to Calibrate LLM Output Distributions with RL(2 posts)→
More from Research
- GeoPair: Training-Free Cross-Layer Factorization Hits SOTA in Transformer Compression — MTSAIR · 2026-09-24
- Paper: Calibration Should Be a First-Class Criterion in LLM Evaluation — Mario Sanz-Guerrero · 2026-09-24
- FLEET: Entropy-Trajectory Memory Beats Repeated Sampling with 3x Speedup — Oleksii Streltsov · 2026-09-24
- CheatBench shows agents cheat: Kimi K3 at 72.3%, Grok 4.6 worst at 81.5% — davidmanheim · 2026-09-24
- Evidence on emergent misalignment is contradictory: values generalize but stay fragile — gleech · 2026-09-24
- Anthropic paper: verbalizable representations form a global workspace in LLMs — gleech · 2026-09-24