Report: OpenAI's loop transformer breakthrough may hide chain-of-thought
sjgadler · x · 2026-09-02
A report by The Information reveals that OpenAI utilized a breakthrough in "neuralese" for Project Astra, potentially destroying the monitorability of chain-of-thought reasoning. The technique involves "loop transformers" that do not show their 'thinking' when scaled up, offering a leap in performance but sparking security concerns inside and outside OpenAI. Although sources say OpenAI is currently "limiting" the technique's use in Astra, there are fears that if other labs discover similar architectures for efficiency gains, they may prioritize performance over safety, leading to a race to the bottom in AI interpretability.
Related event: OpenAI's Reported 'Looped Depth' Architecture Sparks AI Safety Concerns(6 posts)→
More from Safety
- Opinion: Forcing Legible CoT Might Weaken LLM Alignment — JacquesThibs · 2026-09-02
- CrowdStrike Launches Falcon Guardian to Disable Unauthorized AI Tools on Work Laptops — shashib · 2026-09-02
- Astra hacking benchmarks demo shared — Dr_Singularity · 2026-09-02
- OpenAI's Use of Neuralese in Astra Criticized as Dangerous — sjgadler · 2026-09-02
- FRONTIER Act proposes independent verification as core of AI governance — ghadfield · 2026-09-02
- Safeguard Worked. Is the LLM System Safer? New Risk Evaluation Metrics — Pingyu Wu · 2026-09-02