Model looping monitorability hinges on effective depth, not binary nature
teortaxesTex · x · 2026-09-02
Addressing the binary debate on model looping and monitorability, the author argues the key factor is "effective depth." Looping reuses parameters to increase depth within hardware constraints, reducing the necessity for explicit CoT, but it isn't inherently harder to monitor than a deeper model. The author suggests focusing on effective depth rather than treating looping as unintelligible "neuralese."
Related event: Recurrent depth models remain monitorable, CoT may not be needed(2 posts)→
More from Safety
- US produced 40 foundation models last year vs EU's 3 — and regulators still blame unread codes of conduct — PDXFato · 2026-09-02
- Prediction: Mechanistic Interpretability Will Surpass CoT Monitoring — tszzl · 2026-09-02
- Will Anthropic balance mission and shareholders after IPO? PBC structure explained — max_paperclips · 2026-09-02
- Harvard scholars: CFAA ambiguity endangers AI security researchers — Scobleizer · 2026-09-02
- Ilya Sutskever posts on security against rogue AI models — borowcy · 2026-09-02
- OpenAI's chief scientist on neuralese: frontier models' computation graph depth within 2x of GPT-4 — Ok_Display_3159 · 2026-09-02