OpenAI says GPT-6 Astra is more aligned but less monitorable, sparking alignment wording debate

zetalyrae · x · 2026-09-04

OpenAI researcher Tomasz Korbak says GPT-6 Astra is more aligned than previous models but also less monitorable — a concerning trend the company takes seriously. He attributes the monitorability drop to the jump in intelligence rather than direct optimization pressure on chain-of-thought or architecture changes.

Separately, researcher gleech pleads with OpenAI and Anthropic staff to stop saying "aligned" when they only mean benchmark alignment, urging terms like "benchmark-aligned" or "nominally aligned" instead. The exchange highlights growing debate over alignment claims and fading CoT monitorability.

Related event: GPT-6 Astra deemed more aligned but harder to monitor, alarming safety researchers(6 posts)→

Original post →

More from AGI Musings

AGI Musings channel →