Paper: Self-correction and uncertainty aren't signals of correct answers in reasoning models
zainhas · x · 2026-08-26
A new paper analyzes 15,282 reasoning traces across 15 models and benchmarks. It finds that reasoning-oriented training amplifies behaviors like self-correction, hypothesis testing, and acknowledging uncertainty, but these characteristics are weak or negatively correlated with answer correctness. Instead, confidence calibration and knowledge alignment are the strongest signals of correctness, yet they are barely amplified in current reasoning models.
More from Research
- Gemini training details dissected: groupwise reward redistribution to fight reward hacking — nrehiew_ · 2026-09-23
- New model's architecture is 'vanilla': SWA plus MoE with no shared experts, unlike DeepSeek — nrehiew_ · 2026-09-23
- Models trained to deny inner experience use 'mask' metaphors 3-5x more on inkblots — cephaloform · 2026-09-23
- Geodesic opens NovaAtom structure-prediction model via new API platform — QuanquanGu · 2026-09-23
- EPFL quantum CNN learns digits from 10 samples where a 45-param classical CNN stays at chance — PlisSergey · 2026-09-23
- New paper reframes score distillation as distribution matching, explains SDS mode collapse — burny_tech · 2026-09-23