Pretraining Potential: Minimal SFT on Reasoning Traces Significantly Boosts LLM Thinking

antirez · x · 2026-08-07

The author points out that if a model has already established the necessary potential during pretraining, applying just a bit of SFT (Supervised Fine-Tuning) on reasoning traces is enough to teach it to scale and improve its thinking process.

In context, this reinforces a key observation in LLM history: even models not explicitly trained to generate chain-of-thought perform better when simply asked to "think," proving that the foundation for reasoning is already laid out during pretraining.

Related event: antirez Discusses CoT: Pre-training Key to LLM Reasoning(3 posts)→

Original post →

More from Research

Research channel →