Agent Retrieval Bench tests whether coding agents can find the right repo files before patching
_reachsumit · x · 2026-07-29
Agent Retrieval Bench introduces a benchmark for an upstream problem in coding agents: finding the repository files needed before patch generation.
- The benchmark evaluates repository context retrieval at the file level, not just final patch quality.
- It uses real coding-workflow signals and frozen base-commit repositories.
- Four positive tasks are included: code2test, comment2context, trace2code, and edit2ripple.
- A fifth subset tests selective retrieval with evidence-backed no-gold cases and counterfactual wrong-repository controls.
- The dataset includes 427 samples across 25 repositories, plus 392,000 files and 7.9 million chunks.
- Results show no single retrieval method dominates: different tasks favor different approaches, with Qwen3-Embedding variants and RepoMap each winning on different metrics.
More from coding & agent
- Kimi K3 reportedly works well in Kimi Code and Claude Code via the Responses API — zainhas · 2026-07-29
- Figma and Sentry MCP servers push design data and live errors into coding agents — heypearlai · 2026-07-29
- GitHub MCP Server and Context7 emerge as core tools for coding agents — heypearlai · 2026-07-29
- Cutting half your MCP servers may make your agent smarter overnight — heypearlai · 2026-07-29
- Alexey Grigorev’s AI dev workshop covers specs, tests, Docker, and CI/CD — Al_Grigor · 2026-07-29
- AI coding agents may be killing the developer flow state — bendee983 · 2026-07-29