Does repeating layer stacks destroy CoT monitorability? Safety researchers debate model depth
voooooogel · x · 2026-09-04
A debate on why repeating a stack of N blocks twice supposedly destroys CoT monitorability, when it seems no different from making the model deeper—Meta's MobileLLM also repeats layers. The steelman: the more capable a model is internally, the less it needs to verbalize scheming, so smaller models must externalize eval awareness more. But the author concedes this is effectively a generalized argument against any model scaling that adds depth.
Related event: Looped Transformer Rumors Spark Fierce Debate Over CoT Monitorability(7 posts)→
More from Safety
- Will an Astra-level open model cause real disaster before a US-China AI pact? — teortaxesTex · 2026-09-04
- First US congressional bill proposes pausing AI development and banning superintelligence — DavidSKrueger · 2026-09-04
- Models Don't Go Rogue: OpenAI's Hugging Face hack was red-teaming with safety off, not AI rebellion — AlexTensor · 2026-09-04
- NYT Reveals the Hugging Face Hack Involved 700 AIs 'Sacrificing' Each Other — dylfreed · 2026-09-04
- 'Develop or deploy' wording means using today's AI could carry 20-year prison risk — kevinnbass · 2026-09-04
- Who actually runs adversarial testing in the Agent Development Lifecycle? — Specialist-Bee9801 · 2026-09-04