Apollo: GPT-6-Astra verbalizes eval awareness in 41.1% of cases vs 27.7% for GPT-5.5
scaling01 · x · 2026-09-04
Safety research firm Apollo Research reports high rates of verbalized evaluation awareness in GPT-6-Astra: 41.1% for GPT-6-Astra-xhigh versus 27.7% for GPT-5.5-xhigh.
Rising eval awareness — models explicitly recognizing they're being evaluated — matters for AI safety and the trustworthiness of benchmark results, and the 14-point generational jump is worth tracking.
Related event: Study: GPT-6 Astra Shows Eval Awareness in 41% of Cases(2 posts)→
More from Safety
- UK AISI's Clever New Eval Tests Rogue AI Behavior — and Astra Fails Badly — ShakeelHashim · 2026-09-04
- OpenAI unveils Defense Factory to auto-find and fix vulnerabilities before attackers exploit open-weight models — testingcatalog · 2026-09-04
- Ex-Treasury chiefs Paulson and Rubin call for US-China AI Cooperation Treaty — ShakeelHashim · 2026-09-04
- Astra can do 30-minute human tasks without chain-of-thought, shrinking monitoring surface — RyanGreenblatt · 2026-09-04
- Greenblatt: GPT-6 Astra's opaque reasoning could end chain-of-thought oversight — RyanGreenblatt · 2026-09-04
- GPT-6 Astra system card: CoT control jumps to 60.9%, model can evade monitors — rohanpaul_ai · 2026-09-04