OpenAI says GPT-6 Astra is more aligned but less monitorable, sparking alignment wording debate
zetalyrae · x · 2026-09-04
OpenAI researcher Tomasz Korbak says GPT-6 Astra is more aligned than previous models but also less monitorable — a concerning trend the company takes seriously. He attributes the monitorability drop to the jump in intelligence rather than direct optimization pressure on chain-of-thought or architecture changes.
Separately, researcher gleech pleads with OpenAI and Anthropic staff to stop saying "aligned" when they only mean benchmark alignment, urging terms like "benchmark-aligned" or "nominally aligned" instead. The exchange highlights growing debate over alignment claims and fading CoT monitorability.
More from AGI Musings
- Writers love AI drafts, readers hate them: suspected AI content gets skipped and authors punished — birchlse · 2026-09-04
- AI won't make the best lawyers cheaper — it may make them worth more — jkubicki · 2026-09-04
- Andreessen: AI agents meeting crypto is the most important investment theme of the next year — learnedall · 2026-09-04
- Millière vs Mitchell: can the intentional stance make an AI bot a genuine believer? — raphaelmilliere · 2026-09-04
- Millière continues: Odysseus Bot could earn the intentional stance if it pursues persistent goals — raphaelmilliere · 2026-09-04
- Frontier model reasoning shifts from sparse motifs to reasoning-first cognition, observer argues — teortaxesTex · 2026-09-04