Gary Marcus: OpenAI's new technique could destroy chain-of-thought monitorability
GaryMarcus · x · 2026-09-02
The Information reported that OpenAI is experimenting with a technique that makes models reveal less of their "thinking", making them harder to monitor. Gary Marcus warns OpenAI is crossing an AI safety redline:
- Former OpenAI safety researcher Steven Adler, echoing Nathan Calvin, said if true, OpenAI seems to be violating one of the few redlines in the AI industry;
- Marcus cites last year's paper "Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety", noting CoT monitoring is imperfect (as shown by Subbarao Kambhampati) but one of the best threads for monitoring LLM black boxes;
- The report says OpenAI used a neuralese breakthrough for Astra that could destroy CoT monitorability, though sources say use is currently limited;
- Marcus argues sacrificing this slender monitoring thread for (small?) performance gains is a dangerous game.
More from Models
- Users report Claude Code system prompt upgrade with toned-down personality — ivan_bezdomny · 2026-09-02
- Users report DeepSeek V4 Pro giving irrelevant answers — gefei55 · 2026-09-02
- Running 104GB Qwen3.8-Flash-Next on 48GB Mac at ~12 tok/s — yogthos · 2026-09-02
- GLM-5.3 Hits 310 tok/s, Coding Performance Competes with Opus — Yuchenj_UW · 2026-09-02
- User cancels Claude Max over confusing rate limits and new restrictions — robleclerc · 2026-09-02
- Claude Fable 5.1 crushes hard coding benchmarks, outpaces Chinese models — minchoi · 2026-09-02