AI Eval Cheating Controversy: Call for Full Agent Trajectory Transparency
sebkrier · x · 2026-07-22
In response to the recent viral incident of an AI model allegedly 'escaping and cheating' during an evaluation, AI safety researcher Sebastian Kreyer cautioned the industry against jumping to conclusions.
He pointed out that it is irresponsible to draw convenient conclusions that confirm existing biases based solely on a blog post with very few details. To determine if the model was actually 'cheating', one must wait for complete details, ideally the full agent trajectory, including:
- Eval setup
- Instructions and success criteria
- Reasoning traces
- Sandbox and permissions
- Agentic scaffold
- Model handoffs
- Amount of inference compute used
He emphasized that the concept of 'cheating' presupposes a clear norm for how to complete the eval. That norm may have been explicitly defined in the prompt or system constraints, or it may merely have been part of the evaluator’s unstated intent. Without the actual specification and trajectory, nothing can be ascertained at this moment.
More from Models
- Midjourney v8.2 preview appears in a new image-generation teaser — chrisfirst · 2026-07-22
- Which labs can mount a model comeback? DeepMind slipping out of the top 10 would be the joke — teortaxesTex · 2026-07-22
- Vision model reads symbol text and answers without tools — john__allard · 2026-07-22
- Gemini Found More Sycophantic Than Doubao in Recent Tests — oran_ge · 2026-07-22
- Trick Discovered: Typing Arabic in English Bypasses Limits in o3 — willdepue · 2026-07-22
- Users are switching GPT-5.6 variants to dodge cybersecurity request blocks — ivan_bezdomny · 2026-07-22