Microsoft Introduces LoopsBench for Long-Horizon Coding Agents; Top System Solves Only 25% of Tasks

机器之心 · wechat · 2026-08-22

Microsoft, in collaboration with Nanjing University, has released LoopsBench, a benchmark designed to evaluate Coding Agents on long-horizon software engineering tasks. Unlike traditional benchmarks like SWE-bench that focus on one-off issue resolution, LoopsBench emphasizes continuous execution, task dependency management, and regression control.

Key Innovations:

Experimental Findings:

Original post →

More from coding & agent

coding & agent channel →