Does repeating layer stacks really kill CoT monitorability? Debating recurrent architectures vs just deeper models
andersonbcdefg · x · 2026-09-04
In a debate on CoT interpretability, @NathanCalvin argues it's plausible CoT monitoring is long-run doomed yet worth delaying as long as possible. andersonbcdefg pushes back: why does repeating a stack of N blocks twice "magically" destroy CoT monitorability, when it seems no different from making the model deeper — Meta's MobileLLM also repeats layers. Is it just "deep model bad"?
Related event: Looped Transformer Rumors Spark Fierce Debate Over CoT Monitorability(7 posts)→
More from Research
- Alibaba-NLP CORE: boosting compositional reasoning in MLLM embeddings via reranker distillation — Alibaba-NLP · 2026-09-04
- On-policy distillation improves for hundreds of steps from a single query — it's algorithm-starved, not data-starved — Thinking-Space · 2026-09-04
- WorldReward: a vision-language reward model for evaluating camera-conditioned world models — Yibin Wang · 2026-09-04
- Salesforce research: Random eviction of reasoning tokens matches selective KV cache compression — Salesforce · 2026-09-04
- Google Trends quietly redraws samples daily, shrinking significance in 70% of replicated econ papers — RexDouglass · 2026-09-04
- Immunologist builds research-grade flow cytometry software with GPT-6 Astra, cancels all commercial subscriptions — DeryaTR_ · 2026-09-04