Dev Specs 96-Tool MCP Benchmark to Test How Confusable Tool Families Break Agents
EastVersion1226 · reddit · 2026-09-09
An enterprise agent builder proposes a not-yet-run experiment targeting a gap in tool-selection research: papers test thousands of tools from unrelated sources where the right answer is obvious, but real products have 100 sibling tools like getorderstatus vs getorderstatushistory that confuse models.
The experiment spec:
- One coherent 96-tool surface with hand-labelled confusable families, served over a real MCP server, tested on three frontier models
- Key question 1: is degradation caused by the number of choices or the tokens schemas occupy — deciding whether to shard servers or compress schemas
- Key question 2: does performance track total tool count or only within-domain count
- Also measuring spurious tool calls on no-tool queries and error compounding over three turns
The author is asking whether prior work already covers this, and whether production users running 100+ tools see the same problem.
More from coding & agent
- Checkly Rewrote a 92M-Message-a-Day Node Service in Go With AI Agents: Zero Incidents — rseroter · 2026-09-09
- Weaviate Podcast asks whether AutoIndex representation programs can transfer across corpora — CShorten30 · 2026-09-09
- Smart LLM routing cuts costs 69% on 120 tasks while keeping 99.2% success rate — shensi · 2026-09-09
- Test-time adaptation via human-AI interaction paper open-sources full codebase — dan_fried · 2026-09-09
- Users find Google Astra cheaper at higher reasoning levels — brandon_galang · 2026-09-09
- t3code hands-on: manage Claude/Codex/Grok sessions from your phone — kamalgupta09 · 2026-09-09