Paper reframes the objective of LLM RL

The paper “The Mirage of Optimizing Training Policies” argues that LLM reinforcement learning suffers from a mismatch between training and inference policies. It proposes optimizing for a monotonic reasoning policy instead, as a more stable and inference-aligned objective.

2026-07-07 ~ 2026-07-07 · 2 related posts