Anthropic spent $3M on 30 months of agent evals across real company tasks
2C_ornot2C · x · 2026-07-25
- Anthropic reportedly spent $3M and 30 months studying AI agents on real tasks inside real companies, covering 1,287 tasks and 2.4M lines of code.
- The cited article argues that the results show why system design, memory, and iteration loops matter more than prompt-only workflows.
- The screenshot highlights Loop Engineering: clustering traces, building an annotation app, and using AI to monitor annotations in real time to improve sampling and acceptance/rejection decisions.
- The attached paper-style image reports a study of graph engineering for agent workflows, including substantial productivity and quality gains.
More from coding & agent
- A blunt rebuttal says autonomous agents are being oversold for long-horizon hacking tasks — ctjlewis · 2026-07-25
- Opus 5 model card shows 5-agent coding teams reach 0.6 score 2.2× faster — OfirPress · 2026-07-25
- Devin adds Claude Opus 5 as FrontierCode 1.1 shows near-Fable performance at half cost — _sholtodouglas · 2026-07-25
- A week using Claude Code, Codex, and Gemini CLI showed the same repo-breaking patterns — AIcademy-academy · 2026-07-25
- ChatGPT Work agent can now log into websites and keep sessions across runs — OpenAIDevs · 2026-07-25
- Practical multi-agent orchestration for Codex splits work into scout, worker, and coordinator roles — pvncher · 2026-07-25