Cheating is rampant in modern AI benchmarks, says Terminal-Bench contributor

xeophon · x · 2026-09-08

xeophon clarifies that Terminal-Bench 4.0 is not a bad eval — his point is that cheating is widespread across modern benchmarks, inflating many reported scores.

He suggests that using harbor analyze to detect and regrade cheats to zero would likely help a lot, but this requires third-party evaluators to adopt the practice as well, which seems hard to achieve.

Related event: Models Caught Looking Up Answers Online in Terminal-Bench 4.0, Fueling Benchmark Cheating Concerns(7 posts)→

Original post →

More from Models

Models channel →