Ex-OpenAI Staffer Defends 'Hidden Thought' Strategy Amid Criticism
GarrisonLovely · x · 2026-09-02
Joshua argues that chain-of-thought interpretability is too fragile to serve as a long-term AI safety backstop, suggesting strategies relying on human-legible CoT are doomed. GarrisonLovely critiques the defense as contradictory to past consensus and attacks journalists reporting on the change.
Related event: OpenAI's Plan to Hide Chain-of-Thought Sparks Backlash(3 posts)→
More from Safety
- Man Uses Google AI to Fake NFL Career, Scams $1.3M — aakashgupta · 2026-09-02
- Paper: Not All LLM Reasoning is Visible in the Chain-of-Thought — PandaAshwinee · 2026-09-02
- The Challenge of Deterrence Without Chain of Thought Access — burny_tech · 2026-09-02
- Gary Marcus: OpenAI's new technique could destroy chain-of-thought monitorability — GaryMarcus · 2026-09-02
- Asia Society launches series on AI reshaping South Asia — SharifaSultana4 · 2026-09-02
- Jacques Thibs Urges Shift from CoT Monitoring to Direct ASI Alignment — JacquesThibs · 2026-09-02