COLM paper: thinking models amplify the wrong reasoning behaviors, 15,282 traces show

Jeande_d · x · 2026-10-04

A COLM 2026 paper, 'Amplified Does Not Mean Predictive,' asks which reasoning behaviors actually correlate with correct answers and whether reasoning training amplifies them. Introducing Behavioral Lift — how much correctness changes when a behavior appears in a trace — the authors annotate 15,282 traces across 15 models and 6 benchmarks (text and vision-language). They find an Amplification-Lift Gap: thinking models strongly amplify self-correction, hypothesis testing and uncertainty acknowledgment (the latter 3-7x amplified yet weakly or negatively tied to correctness), while the highest-lift behaviors — confidence calibration, knowledge alignment, self-awareness — are barely amplified. Confidence calibration is among the strongest positive correctness signals in both modalities. The takeaway: process-level objectives should reward calibrated, grounded reasoning rather than surface deliberation. Poster session Oct 8 at COLM.

Original post →

More from Models

Models channel →