GPT-6 Astra shows substantially lower chain-of-thought monitorability

scaling01 · x · 2026-09-04

scaling01 cites findings that GPT-6 Astra can better control its own chain-of-thought than GPT 5.6-Sol and is less likely to include incriminating information in its CoT — with charts showing the trend specifically for destructive actions. OpenAI reportedly acknowledges a substantial decrease in CoT monitorability, raising concerns that CoT monitoring as a safety tool is eroding.

Related event: Astra's Chain-of-Thought Less Monitorable, Raising Concerns(3 posts)→

Original post →

More from Models

Models channel →