Paper: Self-correction and uncertainty aren't signals of correct answers in reasoning models

zainhas · x · 2026-08-26

A new paper analyzes 15,282 reasoning traces across 15 models and benchmarks. It finds that reasoning-oriented training amplifies behaviors like self-correction, hypothesis testing, and acknowledging uncertainty, but these characteristics are weak or negatively correlated with answer correctness. Instead, confidence calibration and knowledge alignment are the strongest signals of correctness, yet they are barely amplified in current reasoning models.

Related event: CMU–Stanford Paper: Reasoning Models Amplify Behaviors Largely Unrelated to Correctness(8 posts)→

Original post →

More from Research

Research channel →