Paper reframes the objective of LLM RL
The paper “The Mirage of Optimizing Training Policies” argues that LLM reinforcement learning suffers from a mismatch between training and inference policies. It proposes optimizing for a monotonic reasoning policy instead, as a more stable and inference-aligned objective.
2026-07-07 ~ 2026-07-07 · 2 related posts
- Paper: Train-Inference Mismatch in LLM RL and Monotonic Inference Policies — Jing Liang · 2026-07-06
- 论文:LLM强化学习真正目标是单调推理策略 — _akhaliq · 2026-07-07