Terminal-Bench task accused of contamination: only web-searching models score full marks

xeophon · x · 2026-09-07

X user xeophon flagged that the terminal-bench protein-autointerp-disulfide task has no re-grading, and the only models that solve it are ones that look up the solution online and receive full reward. Among them: GLM-5.3 and Gemini 3.8 Flash. Follow-up data shows Gemini 3.8 Flash uses web access (web search, curl, git clone, HEAD requests) more than all other models combined on tb4.0, raising doubts about how much of its benchmark performance reflects reasoning rather than retrieval.

Related event: Terminal-Bench 4.0 tasks exploited by models searching answers online(5 posts)→

Original post →

More from Models

Models channel →