OpenAI Staff on CoT Monitoring: No Race to Unmonitorability
sjgadler · x · 2026-09-02
Responding to concerns about a race to unmonitorability in AI, an OpenAI employee stated that the company has worked to preserve and utilize chain-of-thought monitoring since its first reasoning models. They emphasized that this technique is vital for seeing how model alignment generalizes. While noting it is fragile and trending negatively, they clarified that the computation graph depth of current frontier models is within a factor of two of GPT-4, and strengthening this monitoring remains a core research goal.
More from Safety
- Scott Alexander: Using anthropomorphism to predict model behavior — repligate · 2026-09-02
- Debate: Is Anthropic intentionally misaligning Claude by prioritizing its 'feelings'? — liminal_bardo · 2026-09-02
- US produced 40 foundation models last year vs EU's 3 — and regulators still blame unread codes of conduct — PDXFato · 2026-09-02
- Prediction: Mechanistic Interpretability Will Surpass CoT Monitoring — tszzl · 2026-09-02
- Will Anthropic balance mission and shareholders after IPO? PBC structure explained — max_paperclips · 2026-09-02
- Harvard scholars: CFAA ambiguity endangers AI security researchers — Scobleizer · 2026-09-02