A Claude Code Suggestion Nearly Broke a Paper's Evaluation
A team withdrew an ICLR submission after discovering that a small fix suggested by Claude Code during OOM debugging had silently corrupted their LLM evaluation, highlighting risks of trusting AI-generated code.
2026-09-26 ~ 2026-09-26 · 2 related posts
- Team withdraws ICLR paper after AI coding agent's fix silently skewed their LLM evaluation — CSProfKGD · 2026-09-26
- Team withdraws ICLR paper after Claude Code's token-budget fix skewed their evals — najoungkim · 2026-09-26