antirez on CoT Essence: A Mix of Sampling Search and State Reasoning

antirez · x · 2026-08-07

Prominent developer antirez shared his fundamental views on the chain-of-thought (CoT) mechanism in LLMs. He believes the model does two things during CoT generation:

Furthermore, he noted that after pretraining, the model already possesses the latent potential to scale and improve its thinking. Just a bit of SFT on reasoning traces is enough for the model to learn how to think better and more deeply.

Related event: antirez Discusses CoT: Pre-training Key to LLM Reasoning(3 posts)→

Original post →

More from Models

Models channel →