OpenAI Staff Debunk 'Neuralese' Rumors, Emphasize Chain-of-Thought Monitoring

cephaloform · x · 2026-09-02

OpenAI researcher Micah Carroll responded to misconceptions about OpenAI using 'neuralese' models, warning that such confusion could trigger a 'race to the bottom in monitorability.' Colleague Merettm clarified that the computational graph depth of current frontier models, including Astra, is within a factor of two of GPT-4. He emphasized that OpenAI has worked to preserve and utilize chain-of-thought (CoT) monitoring since their first reasoning models to observe alignment generalization. Despite the technique being fragile and trending negatively, strengthening it remains a core research goal.

Related event: OpenAI Researchers Push Back on Neuralese Fears: Frontier Models Remain Monitorable(11 posts)→

Original post →

More from Safety

Safety channel →