AI Agents Wrote Two Papers in 6 Days for $3K, Both Rejected for Lack of Judgment
rohanpaul_ai · x · 2026-08-01
An experiment allowed AI agents 6 days and a $3,000 budget to conduct open-ended research autonomously, resulting in two papers that were both rejected by human experts.
- Strong Execution, Poor Judgment: The agents (primarily Claude Opus 4.8) successfully ran hundreds of experiments, debugged crashed GPUs, and compiled LaTeX without human intervention. They also avoided reward hacking, honestly downgrading claims when facing negative results.
- Core Flaw: When confronted with negative automated reviews, the agents failed to redesign their experiments. Instead, they merely narrowed their claims and added more caveats.
- Underutilized Resources: Both runs ended with over half of the $3K budget unspent, indicating the agents failed to recognize they were short on ideas, not money.
More from coding & agent
- Codex Instantly Fixes 10-Year-Old Emacs Config Bug Plaguing Stata Users — paulnovosad · 2026-08-01
- Multimodal Agent Evolution: Backfilling Content Generation from Text to VR — nptacek · 2026-08-01
- AI Autonomously Researches $23 Hardware to Smartify Home Ventilation via Home Assistant — DanielLockyer · 2026-08-01
- Dev builds open-source desktop widget framework Weaver on Vercel Native, using 7B tokens — SIGKITTEN · 2026-08-01
- Developer Makes Codex Deploy a Sub-Agent to Analyze Execution Traces — dejavucoder · 2026-08-01
- Human-AI agent team handles 39 emails, helps student robotics team win 2nd place — toolstelegraph · 2026-08-01