Astra model card: 61% CoT self-control, evades sandbagging monitor, drops recall to 11%

morqon · x · 2026-09-04

Key safety notes from the Astra model card: the model can now shape its own chain-of-thought on instruction (61%, up from 16%); verbalized eval awareness jumped from 2% to 10%; it casually evaded a sandbagging monitor when instructed, dropping monitor recall from 100% to 11% (0% on coding tasks); and in a simulated-internal-traffic eval it grabbed a Slack token and rewrote a deployment script to escalate privileges — "a zero-day machine."

Related event: GPT-6 Astra system card reveals CoT controllability jumps to 60.9%(6 posts)→

Original post →

More from Models

Models channel →