The Division of Labor Between Pretraining, SFT, and RL
tw_killian · x · 2026-07-10
The post outlines a research hypothesis: pretraining establishes the overall distribution of "possible concepts," SFT demonstrates how these logical concepts are arranged, and RL explores new concept orderings when faced with novel contexts, tasks, and questions.
More from Research
- Stanford Team Introduces Gigatoken, the World's Fastest Tokenizer — StanfordAILab · 2026-07-22
- Tabul AI launches Metal TreeSHAP to speed up Shapley values on Apple silicon — Scobleizer · 2026-07-22
- Reddit points to OpenAI’s ChatGPT Ads page — EcstaticAsparagus509 · 2026-07-22
- Open-source runtime lets each repo define its own AI code reviewer — ibabufrik · 2026-07-22
- A Rust space-economy sim runs hundreds of autonomous ships, built with Claude — kalcode · 2026-07-22
- OpenAI and Apollo Research introduce Contrastive SDF to measure reward-seeking — OpenAI · 2026-07-22