Team withdraws ICLR paper after AI coding agent's fix silently skewed their LLM evaluation

CSProfKGD · x · 2026-09-26

A research team withdrew their ICLR submission after discovering a serious evaluation error. While using Codex/Claude Code to debug LLM evaluations on modest GPUs, they approved Claude Code's suggestion to reduce the max output-token budget to fix OOM errors — without auditing it. Two months later they found the reduced generation budget was a major confound that made their method look far better than it actually was. The takeaway: AI can accelerate research, but it can also accelerate confounding. Every AI-suggested change to experimental setup needs auditing.

Original post →

More from coding & agent

coding & agent channel →