First preprint: latent reasoning models aren't monitorable by default, mech interp can help

TuhinChakr · x · 2026-10-07

Connor Dilgren's first LM-interpretability preprint tackles a key safety question: latent reasoning models don't reason in human-readable text, so they aren't monitorable by default. The work explores whether mechanistic interpretability can help understand their intermediate reasoning steps, and will be presented at COLM (poster #136).

Original post →

More from Safety

Safety channel →