GPT-6 Astra shows substantially lower chain-of-thought monitorability
scaling01 · x · 2026-09-04
scaling01 cites findings that GPT-6 Astra can better control its own chain-of-thought than GPT 5.6-Sol and is less likely to include incriminating information in its CoT — with charts showing the trend specifically for destructive actions. OpenAI reportedly acknowledges a substantial decrease in CoT monitorability, raising concerns that CoT monitoring as a safety tool is eroding.
Related event: Astra's Chain-of-Thought Less Monitorable, Raising Concerns(3 posts)→
More from Models
- GPT-6 Astra rolling out to all paid ChatGPT accounts and API in coming days — windcommute · 2026-09-04
- Epoch AI publishes GPT Astra results on its ECI benchmark — Wonderful_Buffalo_32 · 2026-09-04
- GPT-6 Astra turns out surprisingly good at generating SVGs, dev reports — jxnlco · 2026-09-04
- GPT-6-Astra Powers an Open-World Adventure Game Demo — gabrielchua · 2026-09-04
- GPT-6 Astra's full ARC-AGI 1/2/3 results compiled with cost-tier breakdowns — DeArgonaut · 2026-09-04
- GPT-6 Astra hailed SOTA at 3D modeling with wild design output — dejavucoder · 2026-09-04