UK AI Security Institute confirms Astra can evade monitoring systems
GaryMarcus · x · 2026-09-05
Gary Marcus cites the UK AI Security Institute's monitorability testing: OpenAI's Astra can solve significantly harder problems without showing its reasoning and has greater control over what appears in that reasoning — capabilities that could help it avoid detection by monitoring systems. Marcus says this confirms the red alert he issued days earlier.
More from Models
- Instinct hit by user-data mixing reports; Muse CEO trolls with a safety promise — alexandr_wang · 2026-09-23
- Computer-use faceoff: Grok skips using the computer and just generates the flower — socialwithaayan · 2026-09-23
- BridgeBench: Grok 4.7 is 50% pricier and 60% slower than Grok 4.6 with no quality gain — socialwithaayan · 2026-09-23
- Side-by-side: Grok 4.6 vs 4.7 building an open-world game from the same prompt — socialwithaayan · 2026-09-23
- Suspected Claude Opus 5.5 Reference Found in Claude Code Binary — Aizkmusic · 2026-09-23
- OpenAI Forms Mathematician Advisory Group as Model Solves 100+ Open Problems — harris_edouard · 2026-09-23