Thinking intensity does boost capability: R1-Zero paper cited amid DeepSeek max-vs-high setting debate
karminski3 · x · 2026-09-14
Blogger karminski3 weighs in on the debate over whether DeepSeek models should run with max or high thinking intensity, rejecting the claim that reasoning effort doesn't affect performance. He cites the DeepSeek-R1-Zero paper: with pure RL and no new knowledge added, simply letting chains of thought grow longer lifted AIME24 scores from 15% to 71% — evidence that stronger thinking intensity means stronger capability.
He links three discussions: why max settings can invert SWE-bench results, whether max is actually optimal, and how to make models stop thinking at the right time. The linked Fudan NLP paper, Reward Hacking in the Era of Large Models, surveys how optimizing against imperfect proxy signals in RLHF/RLVR breeds reward hacking and emergent misalignment.
His thesis: a true SOTA model should be a perfect information compressor — solving the hardest problems with the fewest tokens, not dumping lengthy reasoning before an answer.
Related event: Debate Over DeepSeek Thinking Intensity: Does Max Mode Really Matter?(2 posts)→
More from Models
- David Bellamy clarifies his experiment used K2 Horizon, an open-weights 375B LLM — JeremyNguyenPhD · 2026-09-14
- Swift-Qwen3.8-27b, a token-efficient reasoning Qwen finetune, trends on Hugging Face — ukisai · 2026-09-14
- GPT-6 Astra hands-on: composes first, orchestrates later, and reportedly outshines Fable and Sol — paw_lean · 2026-09-14
- Researcher posts proof he both synthesized viruses and trained a 375B open-weight LLM — ethanCaballero · 2026-09-14
- Toby Ord: 10x more RLVR compute cuts tokens-to-target ~3x; gains may be math-specific — tobyordoxford · 2026-09-14
- New scaling curve has half the slope: 10,000x compute for 20%-to-80%, but bigger generational jumps — tobyordoxford · 2026-09-14