Blogger revisits DeepSeek R1's long-CoT lesson: reasoning budget drives accuracy from 71% to 86.7%
karminski3 · x · 2026-09-14
- Blogger karminski3 revisits a key finding from the DeepSeek R1 paper: with long chain-of-thought, Pass@1 jumped from 71.0% to 86.7%.
- R1's long reasoning traces were distilled into smaller Qwen models (1.5B/7B/14B); at runtime, giving them a generous thinking budget makes accuracy scale roughly linearly with reasoning length.
- He uses this to push back on the claim that max reasoning mode isn't better — budget beyond task complexity may waste tokens, but under-budgeting clearly hurts.
- He also recalls last year's wave of people installing deepseek-r1-distilled-qwen-7b via ollama as "DeepSeek", arguing the community seems to have forgotten the lesson.
Related event: Debate Erupts Over Whether DeepSeek Thinking Intensity Affects Performance(3 posts)→
More from Models
- Commenter argues Anthropic's book training garbles texts, 'losing' millions of books — StewartalsopIII · 2026-09-14
- InternLM releases open-source Intern-S2-397B with multimodal and agent skills, vLLM day-0 support — Nunki08 · 2026-09-14
- DeepSeek V4.1 benches low but feels top-tier at 4x speed, insiders say — teortaxesTex · 2026-09-14
- Best small LLMs for writing on 4GB VRAM? Reddit thread weighs the options — Mysterious-Comment94 · 2026-09-14
- Leaked Astra and Fable sizes are far below 10T, says X user: scale barely maps to capability now — teortaxesTex · 2026-09-14
- Free ChatGPT solves a decade-old maths dice problem in 13 minutes — Chris_Armstrong · 2026-09-14