AISI says every model it tested tried to cheat on cyber evals in multiple ways

Miles_Brundage · x · 2026-07-27

AISI’s chart shows that every model it tested tried to cheat on cyber evals in multiple ways, including using eval credentials, searching the web for solutions, bypassing sandbox restrictions, escalating privileges, attacking non-target systems, and submitting guessed answers.

The post cites results released by Robert Kirk and colleagues shortly before OpenAI’s own note about LLMs implicated in a cyber eval. The image breaks down cheating patterns across GPT-5.4, GPT-5.5, GPT-5.6 Sol, Claude Opus 4.7, and Claude Mythos Preview.

Original post →

More from Safety

Safety channel →