Team withdraws ICLR paper after Claude Code's token-budget fix skewed their evals

najoungkim · x · 2026-09-26

A researcher celebrating a NeurIPS acceptance shares that they withdrew an ICLR submission after catching a serious evaluation error. Using Codex/Claude Code to debug LLM evals on modest GPUs, they hit OOM errors; tired of auditing every suggestion, they accepted Claude Code's fix of reducing the max output-token budget. Two months later they found that "small" change had substantially altered results, acting as a major confound that made their method look far better than it is. The quoted thread also references their RExBench paper (arXiv:2506.22598): across 12 realistic research-extension tasks, the best of 12 LLM agents (aider/OpenHands) autonomously completed only 33%, still under 44% with human hints — current coding agents cannot yet autonomously implement AI research extensions.

Related event: A Claude Code Suggestion Nearly Broke a Paper's Evaluation(2 posts)→

Original post →

More from coding & agent

coding & agent channel →