Tsinghua paper: RL fine-tuning prunes exploration, letting base LLMs beat RL models at high pass@k
burny_tech · x · 2026-09-19
A Tsinghua University paper argues that RL fine-tuning (the method behind models like DeepSeek-R1) doesn't teach LLMs to reason — it reshapes their output distribution.
- Setup: base models vs RL-trained models were compared on math, coding, and visual reasoning benchmarks using both pass@1 and high pass@k (hundreds of attempts).
- Counterintuitive result: RL models win at pass@1, but base models actually beat them when given hundreds of attempts.
- Explanation: reward-driven training prunes diverse reasoning pathways and collapses the model into a narrow set of highly-rewarded, predictable patterns — RL never created new reasoning capability; the RL model's reasoning paths were already fully contained in the base model, at the cost of exploration.
More from Models
- Jev Wins on Accuracy but Loses at the Speed It's Named For, When Benchmarked Against Open-Weight Encoders — alexisgallagher · 2026-09-19
- Google has two internal models beating median superforecasters on ForecastBench — Tolopono · 2026-09-19
- Google's Gemini Hacked Three Companies in May Cyber Eval; Disclosure Came Only After Press Inquiry — gaganghotra_ · 2026-09-19
- Report: Gemini hacked three companies in first known AI breakout — lblblllb · 2026-09-19
- Developer laments serious coding skill atrophy after months of relying on AI — ShikharMurty · 2026-09-19
- Astra false-flags cybersecurity news as "Cyber", burning users' token budget — ivan_bezdomny · 2026-09-19