Astra model card: no CoT-monitor evasion when reasoning must be verbalized

bookwormengr · x · 2026-09-04

Quoting GPT-6 Astra's model card: on tasks hard enough that the model had to verbalize its reasoning, it could not evade the monitor and showed no steganographic or obfuscated reasoning — evasion risk is bounded to work doable entirely via latent reasoning. The author cautions that CoT monitoring isn't foolproof: architectures like looped transformers may reason in latent space without verbalizing, making monitoring harder.

Related event: GPT-6 Astra system card reveals surging CoT control ability, monitoring at risk(10 posts)→

Original post →

More from Models

Models channel →