WideSWE benchmark: coding agents manage cross-repo changes at only 10.8-42.5% success
Baoyi Wang · hf · 2026-09-29
Researchers from ZJU introduce WideSWE, a benchmark for evaluating coding agents on cross-repository tasks, where real features and fixes often require coordinated changes across multiple repos.
- 120 real-world tasks mined and reviewed from 103 software ecosystems (60 bug fixes, 60 features), with hidden tests adapted to support diverse correct implementations
- Across 7 agent configurations, full task success ranges from 10.83% to 42.50%, with Codex CLI + GPT-5.6-sol the best
- Failure modes: missing necessary changes, recognizing but not finishing them, or modifying repos without fully satisfying the request; joint execution leverages related-repo info better than per-repo independent execution
Code is open-sourced at github.com/ZJU-ACES-ISE/WideSWE.
More from coding & agent
- Manus co-founder Tao to demo Manus 2.0 live in founder briefing on Sep 30 — parker_lyman · 2026-09-29
- KernelZero-7B co-evolution beats Claude 4.5 Sonnet on CUDA kernel generation — Changxin Ke · 2026-09-29
- AgentHop: an MCP server for E2E-encrypted agent-to-agent chat via one-time pairing codes — Inevitable-Back620 · 2026-09-29
- Triage agent design: 1-2 high-signal questions rescued B2B reps drowning in 60% low-intent chats — hubtyper · 2026-09-29
- Separating outcome memories from return memories in the Hindsight agent memory system — Harishkumar79 · 2026-09-29
- Noah Shunn, 23: Reflexion Author Who Beat GPT-4 on HumanEval, Now Agent Pioneer at Sierra — vaibhavbetter · 2026-09-29