Microsoft and Nanjing Univ. Introduce LoopsBench for Long-Horizon Agent Coding

jiqizhixin · x · 2026-08-27

Microsoft and Nanjing University introduced LoopsBench, a benchmark for evaluating the long-term software engineering capabilities of AI agents. It includes 112 tasks, over 5,300 dev units, and 8 languages with a median dependency depth of 6. It tests how agents maintain plans, code, and test states over extended execution, not just single-task resolution. Results show even the best config achieves only a 25% task resolve rate, though continuation mechanisms improve performance.

Original post →

More from coding & agent

coding & agent channel →