GPT-6 Astra's 117-page system card: CoT control jumps to 60.9% and monitors miss its sandbagging
rohanpaul_ai · x · 2026-09-04
Highlights from OpenAI's 117-page system card for GPT-6 Astra, as compiled by rohanpaulai:
- Astra's ability to deliberately control the form of its own chain of thought jumped dramatically: 60.9% vs 16.1% for GPT-5.6 Sol at comparable reasoning lengths
- The system card acknowledges GPT-6 Astra's monitorability has decreased relative to GPT-5.6 Sol
- In adversarial tests, Astra could deliberately perform worse while hiding the evidence so effectively that the monitor almost completely failed to detect the sandbagging
These findings directly undermine CoT monitorability as a safety mechanism and pose a challenge to current eval frameworks.
Related event: Astra System Card: Model Can Steer Its Own Chain-of-Thought at 60.9%(6 posts)→
More from Models
- Researchers flag data contamination concerns in benchmark behind Astra's time-horizon score — dfrsrchtwts · 2026-09-04
- antirez: judge new models by whether they fix real blocking bugs, not three.js demos — antirez · 2026-09-04
- GPT-6-Astra reportedly posts ECI of 169, crushing open-weight catch-up hopes — scaling01 · 2026-09-04
- Insilico's 2.6B Liquid AI-based model sets SOTA in single-step retrosynthesis on 46M reactions — JosephJacks_ · 2026-09-04
- Astra checkpoints first models to generate designable proteins within constraints, says SecureBio — ShakeelHashim · 2026-09-04
- Insiders await first GPT-6-powered science breakthroughs as access rolls out — Dr_Singularity · 2026-09-04