Meta Team Proposes MaD-RL to Calibrate LLM Output Distributions with RL

A former Meta team published MaD-RL, a paper on calibrating LLM output distributions with reinforcement learning. The authors note that reward-maximizing RL like GRPO collapses probability onto specific modes, and correctness-only optimization harms response diversity.

2026-09-24 ~ 2026-09-24 · 2 related posts