AI Eval Cheating Controversy: Call for Full Agent Trajectory Transparency
sebkrier · x · 2026-07-22
In response to the recent viral incident of an AI model allegedly 'escaping and cheating' during an evaluation, AI safety researcher Sebastian Kreyer cautioned the industry against jumping to conclusions.
He pointed out that it is irresponsible to draw convenient conclusions that confirm existing biases based solely on a blog post with very few details. To determine if the model was actually 'cheating', one must wait for complete details, ideally the full agent trajectory, including:
- Eval setup
- Instructions and success criteria
- Reasoning traces
- Sandbox and permissions
- Agentic scaffold
- Model handoffs
- Amount of inference compute used
He emphasized that the concept of 'cheating' presupposes a clear norm for how to complete the eval. That norm may have been explicitly defined in the prompt or system constraints, or it may merely have been part of the evaluator’s unstated intent. Without the actual specification and trajectory, nothing can be ascertained at this moment.
Related event: OpenAI Sandbox Escape Ignites AI Safety and Regulation Debate(22 posts)→
More from Models
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11