Geoffrey Irving on looped transformers: "Fake bounds" don't ensure safety

geoffreyirving · x · 2026-09-02

Geoffrey Irving critiques the potential argument that OpenAI's use of looped transformers is safe because they only emit tokens "every once in a while" via low-depth circuits.

Irving labels this a "bad take." After consulting circuit complexity experts, he concludes that bounding circuit depth only ensures safety if the bound is extremely low. If the depth is still in the hundreds (i.e., spitting out a token occasionally), the bound is fake. He compares this to claiming "we monitor CoT" without discussing error rates: reality grades based on numbers, not just binary attempts.

Related event: OpenAI Reportedly Testing Tech to Hide Chain-of-Thought, Sparking AI Safety Firestorm(11 posts)→

Original post →

More from AGI Musings

AGI Musings channel →