GPT-6 Astra is more aligned but less monitorable — and OpenAI's safety report may not prove containment
sjgadler · x · 2026-09-04
sjgadler amplifies tylertracy321's critique of OpenAI's GPT-6 Astra safety report, quoting researcher tomekkorbak: Astra is more aligned than prior models but less monitorable — a drop OpenAI attributes to an intelligence jump rather than optimization pressure on CoT or architecture. Tracy argues the public cannot conclude OpenAI could stop a rogue deployment: the report suggests Astra could act stealthily internally without timely alerts; internal evals are hard to assess externally; the human escalation process is underexplained; and attack/stealth capabilities may have been under-elicited, risking underestimation.
More from Safety
- System card data contradicts OpenAI's Astra alignment claim, critic says GPT-5.5 safer — GarrisonLovely · 2026-09-04
- Author discovers publisher claimed 100% of Anthropic settlement money for his own books — Elijah_Meeks · 2026-09-04
- repligate argues a post-Pause world would be worse at solving alignment than ours — repligate · 2026-09-04
- UK AISI report fuels doubts that OpenAI's Astra is really distinct from its risky predecessor — GarrisonLovely · 2026-09-04
- US brief calls LLM training fair use, sparking debate on LLM anthropomorphism — TuhinChakr · 2026-09-04
- Anthropic Expands Project Glasswing to 150 Orgs After Partners Found 10,000+ Flaws — ConradBastable · 2026-09-04