Meta Paper: Quantization Causes Overthinking in Reasoning Models, Fixable with Word Penalties

rohanpaul_ai · x · 2026-08-09

A new paper from Meta highlights that post-training quantization in reasoning models introduces noise, causing them to frequently second-guess themselves even after reaching the correct answer.

This quantization noise makes models more likely to generate words like "wait" or "but" at uncertain steps, reopening the reasoning process. Aggressive quantization was found to increase overthinking failures by up to 52%.

Tested across math, coding, and science tasks on models ranging from 1.5B to 32B parameters, the researchers demonstrated that applying a small decoding penalty to about 50 hesitation words can reduce reasoning length by 12% to 23% while maintaining or even improving accuracy.

Original post →

More from Models

Models channel →