COLM paper: latent reasoning traces decodable 65-93% of the time in LRMs

sarahwiegreffe · x · 2026-10-07

Sarah Wiegreffe's team presents a COLM 2026 study on latent reasoning model (LRM) interpretability. Key findings: latent tokens are often unnecessary (LRMs produce nearly identical answers without them on logical reasoning); when needed, gold reasoning traces can be decoded for 65–93% of correct predictions; a new method recovers verified natural-language traces without gold references. Interpretability itself can serve as a signal of prediction correctness.

Original post →

More from Research

Research channel →