Astra model card: no CoT-monitor evasion when reasoning must be verbalized
bookwormengr · x · 2026-09-04
Quoting GPT-6 Astra's model card: on tasks hard enough that the model had to verbalize its reasoning, it could not evade the monitor and showed no steganographic or obfuscated reasoning — evasion risk is bounded to work doable entirely via latent reasoning. The author cautions that CoT monitoring isn't foolproof: architectures like looped transformers may reason in latent space without verbalizing, making monitoring harder.
More from Models
- Ethan Mollick: Treat Fable and Astra class models like an outside team, not an intern — emollick · 2026-09-04
- User observes newer model's safety classifiers appear far more lenient, suspects thoughtcrime training — repligate · 2026-09-04
- Gemini Live gains Google Workspace connectors for Gmail, Keep and Docs — testingcatalog · 2026-09-04
- As GPT-6 Astra Launches, Redditors Recall Their Personal "THE Moment" With AI — Sharp_Caregiver_1534 · 2026-09-04
- 'Astra proves how wrong I was': insider revises his skepticism on AI computer use — sandersted · 2026-09-04
- Unverified: OpenAI said to launch GPT-6 Astra, trained on 100k GPUs, claimed as AGI — AI寒武纪 · 2026-09-04