Study Reveals LLM 'Overthinking': More Reasoning Tokens Can Lead to Wrong Answers

_jaydeepkarale · x · 2026-08-05

Researchers testing OpenAI's open-weight models documented a counterintuitive pattern: on some questions, models consumed more reasoning tokens yet still arrived at incorrect answers.

This phenomenon, termed 'overthinking', occurs because each reasoning token conditions on previous ones. If the model starts down a flawed reasoning path, it spends hundreds of tokens elaborating and reinforcing the mistake rather than catching it. Consequently, defaulting to high reasoning effort on every call does not guarantee better accuracy and significantly inflates costs.

Related event: Study Finds LLM Self-Reflection Ineffective or Even Harmful(3 posts)→

Original post →

More from Research

Research channel →