Anthropic spent $3M on 30 months of agent evals across real company tasks
2C_ornot2C · x · 2026-07-25
- Anthropic reportedly spent $3M and 30 months studying AI agents on real tasks inside real companies, covering 1,287 tasks and 2.4M lines of code.
- The cited article argues that the results show why system design, memory, and iteration loops matter more than prompt-only workflows.
- The screenshot highlights Loop Engineering: clustering traces, building an annotation app, and using AI to monitor annotations in real time to improve sampling and acceptance/rejection decisions.
- The attached paper-style image reports a study of graph engineering for agent workflows, including substantial productivity and quality gains.
More from coding & agent
- Cheaper OpenAI Agents API alternative: sandbox service undercutting E2B by 46% — airesearch12 · 2026-09-11
- His agent kill switch ran for months before he found it was wired to nothing — AnvilandCode · 2026-09-11
- Kernel's Browser Agents Can Now Pay Online Using Aliases, Never Touching Card Data — jeff_weinstein · 2026-09-11
- OpenAI opens up agent sandboxes: BYO or pick from Cloudflare, E2B, Modal, Vercel and more — threepointone · 2026-09-11
- SocialCrawl MCP lets agents search Reddit, YouTube, TikTok, X with one API key — dooddyman · 2026-09-11
- Astra builds a surprisingly polished Catan game in three.js, reusing past UI and 3D assets — FinanceYF5 · 2026-09-11