SpIDER paper: graph-guided code retrieval helps coding agents locate buggy code
mangahomanga · x · 2026-10-11
An EMNLP 2026 main conference paper introduces SpIDER (Spatially Informed Dense Embedding Retrieval), targeting a core bottleneck for LLM coding agents: retrieving relevant code units from large codebases given a bug report or feature request.
- Key idea: Both sparse (BM25) and dense embedding retrievers ignore the graph-structured nature of codebases. SpIDER combines LLM-based reasoning with graph-based exploration, expanding candidates along the repository's code graph on top of an existing retriever.
- Auditable and cheap: The graph is built on-demand from per-repository syntax trees (no offline precomputation); each expanded candidate carries a structural reason (its seed and edge type), keeping the retrieval budget fixed while making the candidate set auditable.
- New benchmark: SpIDER-Bench, a graph-structured benchmark curated from SWEPolyBench, SWE-bench-Verified and Multi-SWE-bench, spanning Python, Java, JavaScript and TypeScript repos.
- Results: SpIDER consistently improves issue localization over existing baselines.
Related event: SpIDER Paper Helps Coding Agents Locate Code Before Editing(2 posts)→
More from coding & agent
- Stop Calling Yourself 'Just a Vibe Coder': Learn Enough That Agents Can't Fool You — brandon_galang · 2026-10-11
- Training inside the harness lifts Qwen3-14B from 22.2% to 54.8% on Spider 2.0-SQLite — omarsar0 · 2026-10-11
- Every Ditched Personal Agents: Dan Shipper's OpenClaw 'Died' and No One Restarted It — every · 2026-10-11
- This founder runs his entire product distribution from an Obsidian vault pointed at Claude Code — EXM7777 · 2026-10-11
- RSIGym pattern: keep agents in lightweight CPU containers, offload training/inference/evals to services — SucceededMind · 2026-10-11
- One settings change turns Claude Code into a multi-model team: Opus plans, Sonnet codes, Haiku searches — chessbuzz · 2026-10-11