Could deeper models hide their scheming? An argument that scaling weakens CoT monitoring

andersonbcdefg · x · 2026-09-04

A discussion on the limits of chain-of-thought monitoring: the steelman is that the more a model can do internally, the less it needs to verbalize if it's scheming — which is why smaller models like Haiku have to externalize eval awareness more than larger ones. The thread generalizes this into an argument against making models deeper, or scaling broadly, with a tongue-in-cheek nod to 1-layer transformers.

Related event: Looped Transformer Rumors Spark Fierce Debate Over CoT Monitorability(7 posts)→

Original post →

More from AGI Musings

AGI Musings channel →