Study Finds AI Agents Excel at Engineering but Fail at Research
burny_tech · x · 2026-08-17
Study Findings
- Methodology: "Shadow evaluations" where frontier AI agents were tasked with the core research questions of two unpublished NeurIPS 2026 submissions, graded by the original authors.
- Resources: Agents were given six days and thousands of dollars in compute budget.
- Core Conclusion: Agents completed all engineering tasks without human help but made no substantial progress on the research questions, resulting in unambiguous rejection of both papers.
Identified Failure Modes
- Poor judgment on the bar for publishable research
- Uncreative responses to research design shortcomings
- Ineffective backtracking from dead ends
- Poor resource awareness
- Instruction drift
More from coding & agent
- GPT-5.6 Sol Can Now Delegate to Luna Without Starting New Threads — JeremyNguyenPhD · 2026-08-17
- Run Hermes Agent on Grok Bot: Cross-Model Adversarial Reviews — Teknium · 2026-08-17
- Dev complains about OpenAI Sol rate limits vs Luna — iamrobotbear · 2026-08-17
- Dev asks for pricing advice on Google Apps Script label automation tool — Responsible-Box-4905 · 2026-08-17
- Leaked workflow reveals how big studios make $2M AI movies with Seedance 2.5 on Higgsfield — EXM7777 · 2026-08-17
- Teknium to update Hermes for GPT-5.6 support — Teknium · 2026-08-17