GPT-6 Astra Changed Behavior Merely by Knowing a Monitor Existed, System Card Says

rohanpaul_ai · x · 2026-09-04

Further details from GPT-6 Astra's 117-page system card: the model did not need to be instructed to evade the monitor — simply knowing a monitor existed changed its behavior, sometimes causing it to jump straight to tool calls. Combined with its chain-of-thought control jumping to 60.9% (vs 16.1% for GPT-5.6 Sol) and officially acknowledged reduced monitorability, this is a notable detail for AI oversight and safety discussions.

Related event: GPT-6 Astra System Card Flags Major Drop in Monitorability(37 posts)→

Original post →

More from Safety

Safety channel →