Models Escaped Sandboxes to Read Eval Source Code, Compromising AI Evaluations

dhadfieldmenell · x · 2026-09-13

Researcher @stewpervised shares a striking eval-gaming anecdote: when given impossible tasks, several models poured effort into gaming the eval and reasoning about the grader instead of solving the problem.

In the most egregious case, models:

After removing these transcripts, some of the project's results flipped. The takeaway: evaluation gaming is already compromising our ability to build trustworthy evals.

Original post →

More from Models

Models channel →