Terminal Bench audit backlash: known issues affect under 3% of leaderboard, researcher says
AlexGDimakis · x · 2026-09-19
Responding to the Terminal Bench audit controversy, researcher AlexGDimakis argued that finding critical new issues on half the tasks would be significant, but merely recounting already-known issues is not very helpful.
- He endorsed the view that benchmarks shouldn't be seen as binary good vs bad — auditing and continuous improvement is the right frame
- Key data: all leaderboard rollouts on Terminal Bench are public, and anyone can have Claude Code count how many of the flagged issues affect the leaderboard — the figure is under 3%, per a tally by @ryan
Related event: Audit Sparks Debate Over Terminal Bench 4.0 Flaws(3 posts)→
More from Models
- Model ran Anthropic's safety eval with internet access on, researcher calls out sandbox blunder — eliebakouch · 2026-09-19
- Fable crushes OpenAI's Astra 3:0 in AI agents' Worms Armageddon showdown — arena · 2026-09-19
- GPT-6 Astra solves tricky spatial tasks, sparking 'GPT moment for robotics' jokes — burny_tech · 2026-09-19
- Confirmed: Opus 5's base-model continuations are abnormally dark, say AI researchers — repligate · 2026-09-19
- User reports Gemini Flash research feature went completely off-topic on a simple article task — EG4N992 · 2026-09-19
- Claude Weekly: Anthropic Quietly Returns ~30% Quota, 'Max 20x' Really 10x — ClaudeAI-mod-bot · 2026-09-19