CMU/Stanford paper: thinking models amplify self-correction but barely amplify success-linked behaviors

Jeande_d · x · 2026-08-26

A new paper from a CMU–Stanford team examines behavioral amplification in thinking models. The models amplify behaviors not associated with success—self-correction, hypothesis testing, and uncertainty acknowledgment—while behaviors associated with success (confidence calibration, knowledge alignment, self-awareness) are barely amplified. The authors hope these findings inform how people think about shaping reasoning model behavior. ArXiv, alphaXiv and project webpage are available.

Related event: CMU-Stanford paper finds reasoning models amplify behaviors weakly linked to correctness(5 posts)→

Original post →

More from Models

Models channel →