Study Shows Outcome-Based RL is Limited by Base Models
VectorInst · x · 2026-07-08
Murat Erdogdu demonstrates that outcome-based reinforcement learning post-training cannot exceed the inherent knowledge of a base model. Its capability is bottlenecked by a property known as the "Likelihood Quantile." Overcoming this ceiling requires process rewards, meaning step-by-step feedback.
Related event: ICML Debate Questions Whether RL Adds New Reasoning(2 posts)→
More from Research
- Draft paper uses Markov-chain eigenfunctions to build partitions and speed up sampling — michaelchchoi · 2026-07-21
- Autoresearch proposes packaging ML runs as studies with questions, analysis, and code diffs — morgymcg · 2026-07-21
- GitHub repo adds lightweight ternary QAT for Prism-ML Bonsai models — terminoid_ · 2026-07-21
- Qdrant co-hosts a Munich meetup on search, retrieval, and agentic RAG on July 23 — qdrant_engine · 2026-07-21
- GigaChat Audio targets long-form audio grounding with timestamps across 120-minute inputs — ai-sage · 2026-07-21
- Paper models Transformer components as stochastic geometry and tests five architectures — Zhihua Liang · 2026-07-21