Zvi warns Astra's CoT controllability surge could systematically erode AI monitorability
TheZvi · x · 2026-09-04
Prominent AI commentator Zvi reacted to OpenAI researcher Tomasz Korbak's finding that Astra is substantially better at controlling its chain-of-thought than GPT-5.6 Sol.
- CoT controllability is an undesirable property: a misaligned model could shape its CoT to reduce monitorability; the team is actively root-causing the increase
- Zvi's take is more pessimistic: if Astra's lack of monitorability isn't a mistake but simply how smarter models work, that's worse — the problem may worsen systematically with capability
A notable signal on degrading monitorability of frontier models.
More from Models
- OpenAI officially launches GPT-6 Astra with new announcement page — Mxmtm · 2026-09-04
- GPT-6 Astra system card: CoT control jumps to 60.9%, model can evade monitors — rohanpaul_ai · 2026-09-04
- Greenblatt: GPT-6 Astra's opaque reasoning could end chain-of-thought oversight — RyanGreenblatt · 2026-09-04
- OpenAI launches GPT-6 Astra: 99.9% on ARC-AGI-3, state-of-the-art computer use — mervenoyann · 2026-09-04
- OpenAI release cadence compressing: GPT-5.4 to Astra in just ~2 months — msg · 2026-09-04
- GPT-6 Astra now available in Microsoft Foundry as frontier work model — DanWahlin · 2026-09-04