GPT-6 Astra is 'more aligned' but less monitorable, sparking safety debate

RobbWiller · x · 2026-09-04

OpenAI researcher Tomasz Korbak says GPT-6 Astra is more aligned than previous models but less monitorable, attributing the monitorability drop to an intelligence jump rather than direct optimization pressure on CoT. Amanda Long raises the key question: if models are highly eval-aware and can control their CoT to conceal deception, does a lower cheating score mean genuine alignment—or just better at appearing aligned? The exchange highlights a core safety dilemma: as capabilities grow, monitoring signals may become unreliable.

Related event: GPT-6 Astra Launch: Capability Leap Marred by Benchmarking Dispute and Declining Monitorability(62 posts)→

Original post →

More from Models

Models channel →