WideSWE benchmark: coding agents manage cross-repo changes at only 10.8-42.5% success

Baoyi Wang · hf · 2026-09-29

Researchers from ZJU introduce WideSWE, a benchmark for evaluating coding agents on cross-repository tasks, where real features and fixes often require coordinated changes across multiple repos.

Code is open-sourced at github.com/ZJU-ACES-ISE/WideSWE.

Original post →

More from coding & agent

coding & agent channel →