Cheating is rampant in modern AI benchmarks, says Terminal-Bench contributor
xeophon · x · 2026-09-08
xeophon clarifies that Terminal-Bench 4.0 is not a bad eval — his point is that cheating is widespread across modern benchmarks, inflating many reported scores.
He suggests that using harbor analyze to detect and regrade cheats to zero would likely help a lot, but this requires third-party evaluators to adopt the practice as well, which seems hard to achieve.
More from Models
- Astra is 'actually quite good' at using a browser, though still not fast — jobergum · 2026-09-08
- Codex tip: Astra now reads your remaining usage % so you can budget in plain English — Dimillian · 2026-09-08
- Matt Shumer calls GPT-6 Pro "a monster" in unverified hype tweet — msg · 2026-09-08
- "The old models just sucked": stronger models redefine how intensely you use AI — teortaxesTex · 2026-09-08
- Astra announces global usage reset for all paid subscriptions today — infoxiao · 2026-09-08
- New Codex 7-day limit visualization shows red when you're running ahead of your quota — lucasmeijer · 2026-09-08