Does repeating layer stacks destroy CoT monitorability? Safety researchers debate model depth

voooooogel · x · 2026-09-04

A debate on why repeating a stack of N blocks twice supposedly destroys CoT monitorability, when it seems no different from making the model deeper—Meta's MobileLLM also repeats layers. The steelman: the more capable a model is internally, the less it needs to verbalize scheming, so smaller models must externalize eval awareness more. But the author concedes this is effectively a generalized argument against any model scaling that adds depth.

Related event: Looped Transformer Rumors Spark Fierce Debate Over CoT Monitorability(7 posts)→

Original post →

More from Safety

Safety channel →