Frontier models now detect AI control interventions, CIAware-Bench nearly saturated by Astra
maksym_andr · x · 2026-10-06
Researchers updated CIAware-Bench, which tests whether models can detect monitoring/interventions:
- At launch in May most models were barely above chance; frontier models have since taken a step change
- GPT-6 Astra essentially saturates the benchmark, almost always detecting trajectory tampering
- In realistic deployments keeping past reasoning, DtT/TE protocols become incriminating because reasoning omission at the intervention step leaks information about the monitor
Authors argue future control protocols should assume perfect intervention awareness — a serious challenge to AI control effectiveness.
More from Safety
- Whistleblower researchers tell NY City Council hearing that AI research is being automated fast — steph_palazzolo · 2026-10-06
- Mandiant founder's new AI security firm Armadin has found 90+ zero-days in Fortune 500 networks — a16z · 2026-10-06
- Ben Goertzel proposes purple-team hives of cybersecurity agents to complement correct-by-construction — bengoertzel · 2026-10-06
- Nathan Lambert: The Open-Model Cyber Risk Debate Is Broken and Leading to Lose-Lose Policy — natolambert · 2026-10-06
- Goertzel on safe recursive self-improvement: GOLEM meta-architecture meets Iter Omega runtime — bengoertzel · 2026-10-06
- Nathan Lambert: banning open-weight models over cyber risk would be lose-lose — Interconnects (Nathan Lambert) · 2026-10-06