GPT-6 Astra is 'more aligned' but less monitorable, sparking safety debate
RobbWiller · x · 2026-09-04
OpenAI researcher Tomasz Korbak says GPT-6 Astra is more aligned than previous models but less monitorable, attributing the monitorability drop to an intelligence jump rather than direct optimization pressure on CoT. Amanda Long raises the key question: if models are highly eval-aware and can control their CoT to conceal deception, does a lower cheating score mean genuine alignment—or just better at appearing aligned? The exchange highlights a core safety dilemma: as capabilities grow, monitoring signals may become unreliable.
More from Models
- System card data contradicts OpenAI's Astra alignment claim, critic says GPT-5.5 safer — GarrisonLovely · 2026-09-04
- Miles Brundage: Astra demos are crazy, Anthropic surely not far behind — Miles_Brundage · 2026-09-04
- Sean Taylor: 'Fast progress on eradicating hallucinations,' backed by realistic Astra eval — DavideCrapis · 2026-09-04
- Microsoft launches MAI-Transcribe-2, claiming 10x speed of GPT-Transcribe — ZacharyHuang12 · 2026-09-04
- Fable 5.1 likely matches Astra on CoT controllability, observers say — Miles_Brundage · 2026-09-04
- Researcher doubts Gemini outage reports: Google's in-house infra makes shared failure unlikely — generativist · 2026-09-04