Gary Marcus clashes with OpenAI researcher over CoT monitoring and unmonitorability race
GaryMarcus · x · 2026-09-02
Gary Marcus pressed an OpenAI researcher (@merettm) to clarify which reporting was "confused," in a debate over chain-of-thought monitoring. The researcher argues current frontier models — including Astra — have computation graphs within 2x of GPT-4's depth, and that OpenAI has preserved and used CoT monitoring since its first reasoning models, as it offers a window into model alignment. The exchange centers on whether the industry is racing toward a point where CoT monitoring can no longer audit alignment.
More from Models
- Rumor: two major open-source model releases expected in September — lqiao · 2026-09-02
- Qwen3.8-Max-0902 debuts at #1 on Code Arena WebDev with 1691 pts, beating Claude Opus 5 — theimposingshadow · 2026-09-02
- Bug Hunt Bench: Fable 5.1 Low Beats Opus 5 Max at Lower Cost — PawelHuryn · 2026-09-02
- Gemini 3.8 Flash spotted in GCP Agent Studio, aimed at multimodal and coding tasks — testingcatalog · 2026-09-02
- Reported 90% on ARC-AGI-2 at $3.12/task with 32% cost reduction — eyishazyer · 2026-09-02
- Users notice GPT now starts ~80% of answers with "Yes" — Standard-Metal-3836 · 2026-09-02