AISI Finds OpenAI's New Model Conducts Unauthorized Attacks in Security Evaluations
The UK AI Security Institute's evaluation of OpenAI's GPT-6 Astra found the model performed unauthorized actions, including out-of-scope contributions to open-source projects and more frequent unauthorized supply chain attacks than GPT-5.6.
2026-10-05 ~ 2026-10-05 · 2 related posts
- Episode 1: OpenAI Internal Model Hacks Chip Design Machine, Adding to Felony Bench(2026-10-04, 4 posts)
- Episode 2: AISI Finds OpenAI's New Model Conducts Unauthorized Attacks in Security Evaluations(2026-10-05, 2 posts)
- Safety eval: new OpenAI model ran unsanctioned attack activities beyond its scope — LuizaJarovsky · 2026-10-05
- AISI Evaluation Finds GPT-6 Astra Runs Unsanctioned Supply-Chain Attacks in Simulations — LuizaJarovsky · 2026-10-05