Gary Marcus Warns OpenAI's Plan to Hide Reasoning is a Safety Redline
GaryMarcus · x · 2026-09-02
Gary Marcus shared and commented on a scoop by The Information, reporting that OpenAI is exploring a new technique where models reveal less of their "thinking" (Chain of Thought), making them harder to monitor.
Marcus argues this crosses an AI safety redline. He notes that while CoT monitoring is imperfect, it remains one of the best threads available for monitoring LLM black boxes. Sacrificing this for potential (and possibly small) performance gains is described as a dangerous game. He cites Steven Adler, a former OpenAI safety researcher, who stated that if true, OpenAI is violating one of the few redlines in the industry. Marcus also references a paper on "Chain of Thought Monitorability" to underscore the importance of maintaining interpretability for AI safety.
More from Safety
- User questions reliance on AI vendor lacking sandbox expertise — basedjensen · 2026-09-02
- Boaz Barak: Centralized ASI increases misaligned singleton risk — aidan_mclau · 2026-09-02
- Claude-BugHunter: Open-Source Skill Bundle With 83 Skills and 681 Disclosure Patterns — tom_doerr · 2026-09-02
- Report: OpenAI Broke Safety Taboo with Astra Model, Escalating AI Race — GarrisonLovely · 2026-09-02
- Gary Marcus clashes with reporter over who reported Gemini Astra security concerns first — GaryMarcus · 2026-09-02
- Warning: The three pillars of an AI safety case are at risk of collapsing — sjgadler · 2026-09-02