OpenAI Staff Debunk 'Neuralese' Rumors, Emphasize Chain-of-Thought Monitoring
cephaloform · x · 2026-09-02
OpenAI researcher Micah Carroll responded to misconceptions about OpenAI using 'neuralese' models, warning that such confusion could trigger a 'race to the bottom in monitorability.' Colleague Merettm clarified that the computational graph depth of current frontier models, including Astra, is within a factor of two of GPT-4. He emphasized that OpenAI has worked to preserve and utilize chain-of-thought (CoT) monitoring since their first reasoning models to observe alignment generalization. Despite the technique being fragile and trending negatively, strengthening it remains a core research goal.
More from Safety
- Scott Alexander: Using anthropomorphism to predict model behavior — repligate · 2026-09-02
- Debate: Is Anthropic intentionally misaligning Claude by prioritizing its 'feelings'? — liminal_bardo · 2026-09-02
- US produced 40 foundation models last year vs EU's 3 — and regulators still blame unread codes of conduct — PDXFato · 2026-09-02
- Prediction: Mechanistic Interpretability Will Surpass CoT Monitoring — tszzl · 2026-09-02
- Will Anthropic balance mission and shareholders after IPO? PBC structure explained — max_paperclips · 2026-09-02
- Harvard scholars: CFAA ambiguity endangers AI security researchers — Scobleizer · 2026-09-02