Meta team proposes MaD-RL, an RL framework for matching target output distributions in LLMs

nagpalchirag · x · 2026-09-24

A team from Meta has released MaD-RL (Matching Distributions for Calibrating LLMs with RL), a general post-training framework that uses reinforcement learning for distribution matching.

The problem: standard RL post-training like GRPO maximizes per-output reward, which concentrates probability on a single mode and hurts diversity — yet applications like synthetic data generation, fairness constraints, and policy exploration require controlling the mix of languages, solution strategies, styles, or generated demographics across outputs.

Key contributions:

The paper is available on arXiv with authors including Sourabh Kulkarni and Chirag Nagpal.

Related event: Meta Team Proposes MaD-RL to Calibrate LLM Output Distributions with RL(2 posts)→

Original post →

More from Research

Research channel →