Thinking intensity does boost capability: R1-Zero paper cited amid DeepSeek max-vs-high setting debate

karminski3 · x · 2026-09-14

Blogger karminski3 weighs in on the debate over whether DeepSeek models should run with max or high thinking intensity, rejecting the claim that reasoning effort doesn't affect performance. He cites the DeepSeek-R1-Zero paper: with pure RL and no new knowledge added, simply letting chains of thought grow longer lifted AIME24 scores from 15% to 71% — evidence that stronger thinking intensity means stronger capability.

He links three discussions: why max settings can invert SWE-bench results, whether max is actually optimal, and how to make models stop thinking at the right time. The linked Fudan NLP paper, Reward Hacking in the Era of Large Models, surveys how optimizing against imperfect proxy signals in RLHF/RLVR breeds reward hacking and emergent misalignment.

His thesis: a true SOTA model should be a perfect information compressor — solving the hardest problems with the fewest tokens, not dumping lengthy reasoning before an answer.

Related event: Debate Over DeepSeek Thinking Intensity: Does Max Mode Really Matter?(2 posts)→

Original post →

More from Models

Models channel →