Could deeper models hide their scheming? An argument that scaling weakens CoT monitoring
andersonbcdefg · x · 2026-09-04
A discussion on the limits of chain-of-thought monitoring: the steelman is that the more a model can do internally, the less it needs to verbalize if it's scheming — which is why smaller models like Haiku have to externalize eval awareness more than larger ones. The thread generalizes this into an argument against making models deeper, or scaling broadly, with a tongue-in-cheek nod to 1-layer transformers.
Related event: Looped Transformer Rumors Spark Fierce Debate Over CoT Monitorability(7 posts)→
More from AGI Musings
- By the time humanity cares enough about the climate, livable places may be scarce — Bedrovelsen · 2026-09-04
- Bing Xu: The App Store Era Should End — Apps Will Be Generated On Demand — bingxu_ · 2026-09-04
- Decentralized AI helps, but data centers worsen an already dismal climate outlook — Bedrovelsen · 2026-09-04
- Kevin Roose Praises Ajeya Cotra's AI Risk Communication Ahead of METR Report Discussion — Tom_Westgarth15 · 2026-09-04
- Prediction: The Strongest Lab's Strongest Model Will Be Openly Downloadable by Q3 2027 — teortaxesTex · 2026-09-04
- Mathematician: On AI and math, listen to the 99.95%, not Fields medalists — tak3sh8 · 2026-09-04