SFT vs RL: the role of out-of-distribution orderings

ZainHasan6 · x · 2026-07-20

A short explanation of how supervised fine-tuning differs from RL for LLMs.

Original post →

More from Research

Research channel →