Microsoft and Nanjing Univ. Introduce LoopsBench for Long-Horizon Agent Coding
jiqizhixin · x · 2026-08-27
Microsoft and Nanjing University introduced LoopsBench, a benchmark for evaluating the long-term software engineering capabilities of AI agents. It includes 112 tasks, over 5,300 dev units, and 8 languages with a median dependency depth of 6. It tests how agents maintain plans, code, and test states over extended execution, not just single-task resolution. Results show even the best config achieves only a 25% task resolve rate, though continuation mechanisms improve performance.
More from coding & agent
- Hugging Face PR adds in-place editing for remote files, "remote storage will never be the same" — lhoestq · 2026-08-27
- Claude Code v2.1.247 adds cost optimization tool and Admin API support — ashwin-ant · 2026-08-27
- Rumor: Anthropic to add task board for managing Claude sub-agents — daniel_mac8 · 2026-08-27
- X Launches Chat API and XDK, Enabling Agent Creation for X Chat — Baconbrix · 2026-08-27
- Talk Release: How DevRel Must Adapt to the Era of AI Agents — heyneighbor · 2026-08-27
- Hermes desktop GUI supports hot-reloadable plugins for customization — NousResearch · 2026-08-27