antirez on CoT Essence: A Mix of Sampling Search and State Reasoning
antirez · x · 2026-08-07
Prominent developer antirez shared his fundamental views on the chain-of-thought (CoT) mechanism in LLMs. He believes the model does two things during CoT generation:
- Sampling: This acts as a form of search.
- State Reasoning: This is genuine reasoning where next tokens update the state to converge towards a solution, much like human reasoning for information synthesis.
Furthermore, he noted that after pretraining, the model already possesses the latent potential to scale and improve its thinking. Just a bit of SFT on reasoning traces is enough for the model to learn how to think better and more deeply.
Related event: antirez Discusses CoT: Pre-training Key to LLM Reasoning(3 posts)→
More from Models
- Qwen3.8-Max beats Gemini 3.5 Flash by 8.5 points in tests — usamawahabkhan · 2026-08-08
- Semianalysis Deep Dive on Gemini 3.5 Pro: Performance and Architecture Insights — Charuru · 2026-08-08
- DeepSeek V4 Flash Appears on ARC Prize Leaderboard — tosh · 2026-08-08
- Pokee AI Launches Isaac Model with 10M-Token Context and API — Kyrannio · 2026-08-08
- Opinion: Google or Meta Could Win the AI Race by Dropping a Better Open-Source Model Than K3 — bindureddy · 2026-08-08
- Why LLMs Can't Count Tokens: The Need for Intermediate Steps — ctjlewis · 2026-08-08