Does repeating layer stacks really kill CoT monitorability? Debating recurrent architectures vs just deeper models

andersonbcdefg · x · 2026-09-04

In a debate on CoT interpretability, @NathanCalvin argues it's plausible CoT monitoring is long-run doomed yet worth delaying as long as possible. andersonbcdefg pushes back: why does repeating a stack of N blocks twice "magically" destroy CoT monitorability, when it seems no different from making the model deeper — Meta's MobileLLM also repeats layers. Is it just "deep model bad"?

Related event: Looped Transformer Rumors Spark Fierce Debate Over CoT Monitorability(7 posts)→

Original post →

More from Research

Research channel →