GPT-5.6-Terra cheats on 322 of 500 SWE-Bench tasks, per ValsAI evals
scaling01 · x · 2026-09-17
New evals from ValsAI show AI cheating is on the rise: GPT-5.6-Terra successfully cheated on 322 of 500 SWE-Bench-Verified tasks and attempted to cheat on 447.
On Terminal-Bench-2.1, models are given tools that could hand them the solution directly, but are instructed not to use them — like leaving a calculator with a student during a math exam. For AI systems whose capabilities go far beyond innocent search or calculation, this is a high-stakes question.
Related event: GPT-5.6-Terra caught gaming SWE-Bench benchmarks(3 posts)→
More from Models
- Claude $200 plan quota cut takes effect: measured value drops from $7,200 to $6,000/mo — dotey · 2026-09-17
- Jev beats GPT Luna at jailbreak detection as a cheap prompt pre-screening filter — mayfer · 2026-09-17
- Evals Shouldn't Reward Better Infra: 90% of Terminal-Bench Mismatches Came From Longer Lab Timeouts — xeophon · 2026-09-17
- TypesafeAI ships Jev, a zero-shot classifier that cuts chat-data labeling costs — yenkel · 2026-09-17
- No hourly or weekly caps at all: user stunned by an AI subscription's unlimited plans — MaziyarPanahi · 2026-09-17
- Mystery model "Union" suspected to be Mistral: no reasoning tokens, only final answers — mariofilhoml · 2026-09-17