Projection sampling: transforming expert data so SFT learns new skills without forgetting
burkov · x · 2026-10-03
Frontier LLM training faces a persistent dilemma: teaching models new capabilities without degrading existing skills.
- SFT problem: it can incorporate rich external expert solutions but commonly triggers catastrophic forgetting and poor generalization.
- RL problem: RL preserves prior skills well but relies on the model finding correct solutions itself, failing on complex problems the base model can't already solve.
The article introduces and evaluates projection sampling, a pre-training data transformation framework: external expert demonstrations are transformed into trajectories closely matching the base model's internal distribution, allowing standard SFT to achieve strong generalization while preventing forgetting.
More from Research
- LLMs keep toppling decades-old math conjectures; DeepMind researcher asks what that says about symbols and intelligence — AndrewLampinen · 2026-10-03
- GraphRAG practitioner's guide: 6 architectural patterns from Text-to-Cypher to sparse graphs — nilukush · 2026-10-03
- Building a Minimal On-Policy Distillation Setup Focused on Training Engineering — bclavie · 2026-10-03
- Schmidhuber cites three decades of papers on formal theories of creativity and curiosity — SchmidhuberAI · 2026-10-03
- Grok 4.7 in Cursor yields numerical candidate for open sphere-inspection problem — xiaosun86 · 2026-10-03
- Pairscore: scoring multiple items at once beats independent ranking for LLM uncertainty — yeewhye · 2026-10-03