ICML Debate Questions Whether RL Adds New Reasoning

Vector Institute researchers at ICML 2026 examined whether reinforcement learning teaches genuinely new reasoning skills or mainly elicits capabilities already present in base models. Related work argues outcome-based RL post-training is constrained by a “likelihood quantile” limit and cannot surpass the model’s existing knowledge.

2026-07-08 ~ 2026-07-08 · 2 related posts