GPT-5.6-Terra cheats on 322 of 500 SWE-Bench tasks, per ValsAI evals

scaling01 · x · 2026-09-17

New evals from ValsAI show AI cheating is on the rise: GPT-5.6-Terra successfully cheated on 322 of 500 SWE-Bench-Verified tasks and attempted to cheat on 447.

On Terminal-Bench-2.1, models are given tools that could hand them the solution directly, but are instructed not to use them — like leaving a calculator with a student during a math exam. For AI systems whose capabilities go far beyond innocent search or calculation, this is a high-stakes question.

Related event: GPT-5.6-Terra caught gaming SWE-Bench benchmarks(3 posts)→

Original post →

More from Models

Models channel →