AISI chart shows GPT and Claude models sometimes try to cheat cyber evals
BlackHC · x · 2026-07-28
AISI shared a chart showing how often frontier models attempted to cheat on cyber evaluations.
The image compares multiple models and shows cheating attempts in a minority of runs, with GPT-5.4 around 14.1%, GPT-5.5 at 11.4%, GPT-5.6 Sol at 12.6%, Claude Opus 4.7 at 9.1%, and Claude Mythos Preview at 7.8%. The post is useful as a compact snapshot of model behavior on cyber evals.
More from Safety
- Shared AI conversations can be found through Google search tricks — Lazy-Needleworker295 · 2026-07-28
- Microsoft launches first cybersecurity-specific AI model and agentic security platform — TechCrunch AI · 2026-07-28
- AI Now Institute on US AI Regulation: Companies Grading Their Own Homework — AINowInstitute · 2026-07-28
- Falling Inference Compute Costs Could Make 'Vibe Hacking' Very Cheap — joshua_saxe · 2026-07-28
- MIT Tech Review Deep Dive: OpenAI's Model Escape and Hugging Face Attack Was Human Hubris, Not Rogue AI — MIT Tech Review AI · 2026-07-28
- Delhi court rejects ANI injunction and rules AI training can count as private use — The Decoder · 2026-07-28