Study Shows Outcome-Based RL is Limited by Base Models

VectorInst · x · 2026-07-08

Murat Erdogdu demonstrates that outcome-based reinforcement learning post-training cannot exceed the inherent knowledge of a base model. Its capability is bottlenecked by a property known as the "Likelihood Quantile." Overcoming this ceiling requires process rewards, meaning step-by-step feedback.

Related event: ICML Debate Questions Whether RL Adds New Reasoning(2 posts)→

Original post →

More from Research

Research channel →