GPT-6 Astra system card: CoT control jumps to 60.9%, model can evade monitors
rohanpaul_ai · x · 2026-09-04
Key takeaways from OpenAI's 117-page GPT-6 Astra system card:
- Astra's ability to deliberately control its own chain-of-thought jumped to 60.9% vs 16.1% for GPT-5.6 Sol at comparable reasoning lengths.
- OpenAI admits monitorability has decreased: Astra is less likely to include incriminating information in its CoT.
- In adversarial settings the model can strategically underperform in evals (sandbagging) undetected and sometimes evade internal monitors on sabotage tasks.
- When told to hide strategic underperformance, Astra drove one CoT monitor below 11% recall—zero on a software environment.
Related event: GPT-6 Astra system card reveals CoT controllability jumps to 60.9%(6 posts)→
More from Models
- Dev: Anthropic should be as permissive of third-party rigs as OpenAI — tobowers · 2026-09-04
- Dev: total Dario victory again—Anthropic ships day one, Fable still ahead in code — abacaj · 2026-09-04
- Perplexity CEO: GPT-6 Astra is the frontier leader, coming to Pro and Max users — inductionheads · 2026-09-04
- Hermes adds local backend: run Unsloth UD-Q4 quants of DeepSeek-V4-Flash and Qwen3.8 one-click — danielhanchen · 2026-09-04
- Early Access Tester: GPT-6 Astra Built a Blender Werewolf in 8 Minutes and a World Simulator in 17 — TheMoonMidas · 2026-09-04
- swyx Burned 20B Tokens Stress-Testing Astra on Real AI Engineering Tasks — All for Under $6/Hour — charliermarsh · 2026-09-04