RecToolBench: 1,200+ Task Benchmark Tests Recommender Agents on MCP Tool Orchestration
_reachsumit · x · 2026-09-28
A new arXiv paper introduces RecToolBench, an MCP-based benchmark evaluating tool-using recommender agents under fuzzy user instructions.
Scale and structure:
- Over 1,200 executable tasks across three recommendation domains, 13 MCP servers, and 32 tools;
- Covers single-tool calls, parallel calls, sequential tool chains, and hybrid orchestration;
- Built with a scalable synthesize-fuzzify-judge pipeline, evaluating agent trajectories via rule-based execution checks and rubric-based LLM evaluation.
Key findings: experiments on representative LLMs show syntactically valid tool calls don't guarantee successful recommendations — models struggle with semantic parameter grounding, multi-step evidence integration, and grounded final recommendations, especially as orchestration complexity increases.
More from coding & agent
- Personal agents still need a portable 'ask me first' file alongside MEMORY.md — sujingshen · 2026-09-28
- DIY harness × model benchmark: Claude Code hits 100% while deepseek-v4.1-flash matches 96% in half the time — dh7net · 2026-09-28
- ComfyVault dedupes model files across ComfyUI installs via symlinks — ruashots · 2026-09-28
- Agent failed for weeks on a problem StackOverflow answered instantly — itsOmSarraf_ · 2026-09-28
- Harness-Zero: PKU, Google and HKUST Distill Agent Harnesses into Model Weights — AxSaucedo · 2026-09-28
- PromptSentry: open-source 3-layer proxy blocks prompt injections in under 1ms with local DLP scrubbing — Ok-Negotiation342 · 2026-09-28