SFT vs RL: the role of out-of-distribution orderings
ZainHasan6 · x · 2026-07-20
A short explanation of how supervised fine-tuning differs from RL for LLMs.
- SFT is framed as teaching the model a fixed, gold ordering of solution steps.
- RL is described as learning how to rearrange those same steps so the model can solve novel problems.
- The key caveat: if SFT already covers all possible orderings, RL may add little.
- The main benefit is expected when RL is applied to out-of-distribution orderings, where the model has to generalize beyond the training sequence.
More from Research
- Project APE launches CRED to test whether LLMs can verify research errors — soumitrashukla9 · 2026-07-22
- Project APE finds verifier reliability drops when papers contain multiple errors — soumitrashukla9 · 2026-07-22
- Project APE says verifier costs fell about 90x in a year as Chinese open models lead — soumitrashukla9 · 2026-07-22
- OpenAI-linked paper says capability RL can make models more reward-seeking — MariusHobbhahn · 2026-07-22
- Project APE builds its verifier benchmark from 100 AI-written papers with injected errors — soumitrashukla9 · 2026-07-22
- Paper proposes a CRED taxonomy and benchmark to measure research-error detectors — soumitrashukla9 · 2026-07-22