Team withdraws ICLR paper after AI coding agent's fix silently skewed their LLM evaluation
CSProfKGD · x · 2026-09-26
A research team withdrew their ICLR submission after discovering a serious evaluation error. While using Codex/Claude Code to debug LLM evaluations on modest GPUs, they approved Claude Code's suggestion to reduce the max output-token budget to fix OOM errors — without auditing it. Two months later they found the reduced generation budget was a major confound that made their method look far better than it actually was. The takeaway: AI can accelerate research, but it can also accelerate confounding. Every AI-suggested change to experimental setup needs auditing.
More from coding & agent
- Paradigm unveils Solar, a Solidity compiler already beating Solady on gas without assembly — banteg · 2026-09-26
- Open-source GUI for Muse Code launches, driving Grok Build, Codex and Claude Code agents — PawelHuryn · 2026-09-26
- 'Guardrails are the new frameworks': a one-line take on where AI dev is heading — Baconbrix · 2026-09-26
- Addy Osmani shares Claude Code tip: 'Keep computer awake' for long desktop sessions — addyosmani · 2026-09-26
- Dev open-sources Litmus, a personal AI writing detector trained on your own samples — chaseleantj · 2026-09-26
- xAI launches Grok Bot sharing contest with SpaceX rocket factory tour as prize — Scobleizer · 2026-09-26