Team withdraws ICLR paper after Claude Code's token-budget fix skewed their evals
najoungkim · x · 2026-09-26
A researcher celebrating a NeurIPS acceptance shares that they withdrew an ICLR submission after catching a serious evaluation error. Using Codex/Claude Code to debug LLM evals on modest GPUs, they hit OOM errors; tired of auditing every suggestion, they accepted Claude Code's fix of reducing the max output-token budget. Two months later they found that "small" change had substantially altered results, acting as a major confound that made their method look far better than it is. The quoted thread also references their RExBench paper (arXiv:2506.22598): across 12 realistic research-extension tasks, the best of 12 LLM agents (aider/OpenHands) autonomously completed only 33%, still under 44% with human hints — current coding agents cannot yet autonomously implement AI research extensions.
Related event: A Claude Code Suggestion Nearly Broke a Paper's Evaluation(2 posts)→
More from coding & agent
- Max out your Claude subscription: Opus orchestrates, cheap models execute tickets — EXM7777 · 2026-09-26
- The emerging standards behind AI agents: MCP, A2A, ACP and Open Responses each cover one boundary — bibryam · 2026-09-26
- Codex workshop at FIU gets students hands-on building for an hour — paw_lean · 2026-09-26
- llama.cpp fork dedups repeated prompts losslessly, cutting 108k to 71k tokens in agent loops — Odd_Cauliflower_8004 · 2026-09-26
- FlexViz keeps 100M+ row charts interactive with lazy Polars aggregations and Rust kernels — JeremyCMorgan · 2026-09-26
- Built in an hour with Claude Code: the killer AI apps may be the ones we make for each other — alfred_lua · 2026-09-26