Tsinghua paper: RL fine-tuning prunes exploration, letting base LLMs beat RL models at high pass@k

burny_tech · x · 2026-09-19

A Tsinghua University paper argues that RL fine-tuning (the method behind models like DeepSeek-R1) doesn't teach LLMs to reason — it reshapes their output distribution.

Original post →

More from Models

Models channel →