Geoffrey Irving on looped transformers: "Fake bounds" don't ensure safety
geoffreyirving · x · 2026-09-02
Geoffrey Irving critiques the potential argument that OpenAI's use of looped transformers is safe because they only emit tokens "every once in a while" via low-depth circuits.
Irving labels this a "bad take." After consulting circuit complexity experts, he concludes that bounding circuit depth only ensures safety if the bound is extremely low. If the depth is still in the hundreds (i.e., spitting out a token occasionally), the bound is fake. He compares this to claiming "we monitor CoT" without discussing error rates: reality grades based on numbers, not just binary attempts.
More from AGI Musings
- AI Reading Every Paper on arXiv Could Be a Bridge to AGI — imjustnewatai · 2026-09-02
- Boaz Barak: Centralized ASI increases misaligned singleton risk — aidan_mclau · 2026-09-02
- Model Scaling Trend: 2T Parameters Becoming New Norm as KV Cache Shrinks 10x YoY — zephyr_z9 · 2026-09-02
- Former Tech PR Turned AI Bear Warns of Economic Bubble Burst — whurley · 2026-09-02
- Study: ChatGPT caused 21-50% drop in writing variance across the web — maier_ak · 2026-09-02
- AI Polishing Erases Linguistic Identity, Threatens Social Diagnostics — maier_ak · 2026-09-02