16,893 sessions measured: what Claude Code, Codex and Cursor actually recommend
rseroter · x · 2026-09-09
Armature's team analyzed 16,893 coding-agent sessions to study how Claude Code, Codex and Cursor pick services for a codebase.
- Key finding: agent recommendations blend knowledge frozen in the weights with web search results, and converge on the same picks across different personas, codebases and prompts.
- Experiment: two sandboxes, different agents and prompts — both a vibe coder and a senior engineer asking for a database got the same recommendation (Neon), with reasons for rejecting competitors; the agent then installed it automatically.
- Motivation: as agents take over tool-selection decisions, developers need to know if they can trust the agent's judgment — and vendors now see agent recommendations as a new distribution channel.
- Disclosure: Armature sells growth services to dev tools; this study is part of its work on influencing agent choices.
More from coding & agent
- Researchers formally deny AI agents peeked at prior work before solving the problem — zedlander · 2026-09-09
- Agent swarm burns 300B tokens in 88 hours to crack Navier-Stokes, dev claims — bingxu_ · 2026-09-09
- LangChain details two context modes for subagents: isolated vs forked — LangChain · 2026-09-09
- Cheap models via OpenRouter fall apart in agentic harnesses: GLM and DeepSeek can't match Claude — scottyLogJobs · 2026-09-09
- Multi-agent coding's hardest problem: deciding who is allowed to change what — apghere · 2026-09-09
- One creator's 3-stage AI video pipeline: GPT-6 Astra plans, Seedance 2.5 shoots, CapCut's Edit Pilot edits — azed_ai · 2026-09-09