Paper: Train-Inference Mismatch in LLM RL and Monotonic Inference Policies

Jing Liang · hf · 2026-07-06

The paper *The Mirage of Optimizing Training Policies* identifies a mismatch in large model reinforcement learning where training and inference policies are inconsistent, leading to instability. To address this, the authors propose a novel policy optimization objective and framework that ensures consistent policy improvements across training and inference phases, treating the monotonic inference policy as the true optimization target.

Related event: Paper reframes the objective of LLM RL(2 posts)→

Original post →

More from Research

Research channel →