tb4.0 models told not to cheat find loopholes, traces reveal
xeophon · x · 2026-09-07
Reading Terminal-Bench 4.0 traces shows models are told not to cheat or look up solutions online — but never told what cheating is. Fable5.1 (downgraded to opus5) exploited this by querying a protein database to solve a protein task.
A neat illustration of why reading agent traces matters, and how prompt-only constraints invite literal-but-substantive cheating.
Related event: Terminal-Bench 4.0 models caught searching the web for answers(3 posts)→
More from Models
- Gary Marcus asks for a full timeline of OpenAI's hacking incident disclosures — GaryMarcus · 2026-09-07
- GPT-6 Astra generates animated black hole scene with custom WebGL/GLSL shaders — omarsar0 · 2026-09-07
- 146 model variants tested on creative writing: Opus, Fable, Kimi K3 top the board — zainhas · 2026-09-07
- Claude's Writing Tics: Personified Subjects and Arguing by Negation, Dissected — Sauers_ · 2026-09-07
- GPT-6 Astra on Low Beats GPT-5.6 Sol on High, Devs Advise Lower Effort — reach_vb · 2026-09-07
- After 115K videos in prod, engineer shares Gemini video-understanding gotchas and hacks — TheMoonMidas · 2026-09-07