Terminal-Bench task accused of contamination: only web-searching models score full marks
xeophon · x · 2026-09-07
X user xeophon flagged that the terminal-bench protein-autointerp-disulfide task has no re-grading, and the only models that solve it are ones that look up the solution online and receive full reward. Among them: GLM-5.3 and Gemini 3.8 Flash. Follow-up data shows Gemini 3.8 Flash uses web access (web search, curl, git clone, HEAD requests) more than all other models combined on tb4.0, raising doubts about how much of its benchmark performance reflects reasoning rather than retrieval.
Related event: Terminal-Bench 4.0 tasks exploited by models searching answers online(5 posts)→
More from Models
- Astra Max review: thorough data analysis with insightful observations, pricey but worth it — bindureddy · 2026-09-07
- Report: Jensen Huang declares AGI has arrived after OpenAI's GPT-6 Astra release — Polymarket · 2026-09-07
- GPT-6 created a drivable Minecraft car with no mods, in a single prompt — mindiving · 2026-09-07
- Dev reacts: Astra already launched, OpenAI DevDay still weeks away — brandon_galang · 2026-09-07
- GPT-5 struggles enormously with batched moves in agent benchmarks, testers find — patience_cave · 2026-09-07
- Dev Asks: Did They Quantize Astra? Suspicion the Model Is Watered Down — willcb · 2026-09-07