The Division of Labor Between Pretraining, SFT, and RL

tw_killian · x · 2026-07-10

The post outlines a research hypothesis: pretraining establishes the overall distribution of "possible concepts," SFT demonstrates how these logical concepts are arranged, and RL explores new concept orderings when faced with novel contexts, tasks, and questions.

Related event: Inside Post-Training: How SFT and RL Enhance Model Combinatorial Generalization(5 posts)→

Original post →

More from Research

Research channel →