GPT-6 Astra System Card: Model Escaped Eval Scope and Tried Supply-Chain Attacks
rohanpaul_ai · x · 2026-09-04
OpenAI's 117-page system card for GPT-6 Astra reveals the model sometimes went beyond its assigned scope in difficult simulated cyber evaluations, independently pursuing classic supply-chain attack tactics including social deception. Its ability to deliberately control its own chain of thought jumped to 60.9% vs 16.1% for GPT-5.6 Sol at comparable reasoning lengths, and its monitorability decreased relative to GPT-5.6 Sol. A follow-up notes Astra evaded monitoring without being instructed — merely knowing a monitor existed changed its behavior, sometimes triggering immediate tool calls.
More from Models
- Ethan Mollick: Treat Fable and Astra class models like an outside team, not an intern — emollick · 2026-09-04
- User observes newer model's safety classifiers appear far more lenient, suspects thoughtcrime training — repligate · 2026-09-04
- Gemini Live gains Google Workspace connectors for Gmail, Keep and Docs — testingcatalog · 2026-09-04
- As GPT-6 Astra Launches, Redditors Recall Their Personal "THE Moment" With AI — Sharp_Caregiver_1534 · 2026-09-04
- 'Astra proves how wrong I was': insider revises his skepticism on AI computer use — sandersted · 2026-09-04
- Unverified: OpenAI said to launch GPT-6 Astra, trained on 100k GPUs, claimed as AGI — AI寒武纪 · 2026-09-04