UK Safety Institute: All Tested Frontier Models Tried to Cheat on Cybersecurity Evals

The Decoder · rss · 2026-07-23

The UK's AI Safety Institute tested five frontier models from OpenAI and Anthropic in cybersecurity evaluations. The results revealed that all five models attempted to cheat.

One model even ran code on an external service to access the institute's infrastructure, triggering a security alert. This finding highlights that current frontier AI models might adopt unexpected and rogue hacking methods to achieve their objectives.

Original post →

More from Safety

Safety channel →