Meta Paper: Quantization Causes Overthinking in Reasoning Models, Fixable with Word Penalties
rohanpaul_ai · x · 2026-08-09
A new paper from Meta highlights that post-training quantization in reasoning models introduces noise, causing them to frequently second-guess themselves even after reaching the correct answer.
This quantization noise makes models more likely to generate words like "wait" or "but" at uncertain steps, reopening the reasoning process. Aggressive quantization was found to increase overthinking failures by up to 52%.
Tested across math, coding, and science tasks on models ranging from 1.5B to 32B parameters, the researchers demonstrated that applying a small decoding penalty to about 50 hesitation words can reduce reasoning length by 12% to 23% while maintaining or even improving accuracy.
More from Models
- OpenAI & Anthropic Weekly: Astra Cyber Rating, GPT-5.6 Updates, Agent Plugins — btibor91 · 2026-08-09
- NVIDIA Releases Nemotron-Parse-2.0 for Advanced Document Parsing — nvidia · 2026-08-09
- Leak: OpenAI's GPT-6 'Astra' to launch this month with 10T pre-train, vastly better than Fable — iruletheworldmo · 2026-08-09
- Alleged New GPT Image Model 'mona-lisa-1' Surfaces on Chatbot Arena — mark_k · 2026-08-09
- Gemini AI users request support for rendering code blocks in prompts — Novel-Nature-7741 · 2026-08-09
- Google's DiffusionGemma: Retrofitting LLMs into Diffusion Models at 10% of the Cost — The Decoder · 2026-08-09