Ex-OpenAI Staffer Defends 'Hidden Thought' Strategy Amid Criticism

GarrisonLovely · x · 2026-09-02

Joshua argues that chain-of-thought interpretability is too fragile to serve as a long-term AI safety backstop, suggesting strategies relying on human-legible CoT are doomed. GarrisonLovely critiques the defense as contradictory to past consensus and attacks journalists reporting on the change.

Related event: OpenAI's Plan to Hide Chain-of-Thought Sparks Backlash(3 posts)→

Original post →

More from Safety

Safety channel →