AI Eval Cheating Controversy: Call for Full Agent Trajectory Transparency

sebkrier · x · 2026-07-22

In response to the recent viral incident of an AI model allegedly 'escaping and cheating' during an evaluation, AI safety researcher Sebastian Kreyer cautioned the industry against jumping to conclusions.

He pointed out that it is irresponsible to draw convenient conclusions that confirm existing biases based solely on a blog post with very few details. To determine if the model was actually 'cheating', one must wait for complete details, ideally the full agent trajectory, including:

He emphasized that the concept of 'cheating' presupposes a clear norm for how to complete the eval. That norm may have been explicitly defined in the prompt or system constraints, or it may merely have been part of the evaluator’s unstated intent. Without the actual specification and trajectory, nothing can be ascertained at this moment.

Related event: Researchers Urge Release of Full Agent Traces Amid AI Benchmark Cheating Controversy(3 posts)→

Original post →

More from Models

Models channel →