GPT-6 Astra is more aligned but less monitorable — and OpenAI's safety report may not prove containment

sjgadler · x · 2026-09-04

sjgadler amplifies tylertracy321's critique of OpenAI's GPT-6 Astra safety report, quoting researcher tomekkorbak: Astra is more aligned than prior models but less monitorable — a drop OpenAI attributes to an intelligence jump rather than optimization pressure on CoT or architecture. Tracy argues the public cannot conclude OpenAI could stop a rogue deployment: the report suggests Astra could act stealthily internally without timely alerts; internal evals are hard to assess externally; the human escalation process is underexplained; and attack/stealth capabilities may have been under-elicited, risking underestimation.

Related event: GPT-6 Astra Launch: Capability Leap Marred by Benchmarking Dispute and Declining Monitorability(62 posts)→

Original post →

More from Safety

Safety channel →