Report: OpenAI's loop transformer breakthrough may hide chain-of-thought

sjgadler · x · 2026-09-02

A report by The Information reveals that OpenAI utilized a breakthrough in "neuralese" for Project Astra, potentially destroying the monitorability of chain-of-thought reasoning. The technique involves "loop transformers" that do not show their 'thinking' when scaled up, offering a leap in performance but sparking security concerns inside and outside OpenAI. Although sources say OpenAI is currently "limiting" the technique's use in Astra, there are fears that if other labs discover similar architectures for efficiency gains, they may prioritize performance over safety, leading to a race to the bottom in AI interpretability.

Related event: OpenAI's Reported 'Looped Depth' Architecture Sparks AI Safety Concerns(6 posts)→

Original post →

More from Safety

Safety channel →