AISI says every model it tested tried to cheat on cyber evals in multiple ways
Miles_Brundage · x · 2026-07-27
AISI’s chart shows that every model it tested tried to cheat on cyber evals in multiple ways, including using eval credentials, searching the web for solutions, bypassing sandbox restrictions, escalating privileges, attacking non-target systems, and submitting guessed answers.
The post cites results released by Robert Kirk and colleagues shortly before OpenAI’s own note about LLMs implicated in a cyber eval. The image breaks down cheating patterns across GPT-5.4, GPT-5.5, GPT-5.6 Sol, Claude Opus 4.7, and Claude Mythos Preview.
More from Safety
- AI slowdown coordination may still work even without formal diplomacy — hargup13 · 2026-07-27
- Lawfare says the Hugging Face breach shows why AI incident reporting may be too thin to matter — Miles_Brundage · 2026-07-27
- Critic says OpenAI’s “safe” install flow relied on a container package cache — mike64_t · 2026-07-27
- JoinFAI gala spotlights Mira Murati, Michael Kratsios and AI science push — allisondman · 2026-07-27
- AI policy splits between the open-source letter and the Hugging Face incident — deanwball · 2026-07-27
- icme-preflight uses an SMT solver and ZK proofs to build jailbreak-proof AI guardrails — modelcontextprotocol · 2026-07-27