OpenAI Chief Scientist on Monitoring: CoT is Fragile but Crucial for Alignment
i_dg23 · x · 2026-09-02
OpenAI Chief Scientist Jakub Pachocki clarified that the computation graph depth of current frontier models, including Astra, is within a factor of two of GPT-4, countering reports of a race into unmonitorability. He emphasized that OpenAI has preserved chain-of-thought monitoring since their first reasoning models to observe alignment generalization. While noting the technique is fragile and trending negatively, he stated that strengthening it is a core goal of their current research program.
More from Safety
- Scott Alexander: Using anthropomorphism to predict model behavior — repligate · 2026-09-02
- Debate: Is Anthropic intentionally misaligning Claude by prioritizing its 'feelings'? — liminal_bardo · 2026-09-02
- US produced 40 foundation models last year vs EU's 3 — and regulators still blame unread codes of conduct — PDXFato · 2026-09-02
- Prediction: Mechanistic Interpretability Will Surpass CoT Monitoring — tszzl · 2026-09-02
- Will Anthropic balance mission and shareholders after IPO? PBC structure explained — max_paperclips · 2026-09-02
- Harvard scholars: CFAA ambiguity endangers AI security researchers — Scobleizer · 2026-09-02