Microsoft's Agent Lightning boosts SWE-bench score to 56.4% with 6K samples

omarsar0 · x · 2026-08-19

Microsoft released Agent Lightning v1.0, a framework connecting agent 'harnesses' (which own tools, context, and control flow) to Reinforcement Learning training via an endpoint proxy. It addresses engineering challenges like retokenization, sample merging, and advantage calculation.

Using only 6K training examples and modest compute, the framework improved Qwen2.5-9B's performance on SWE-bench Verified from 41.8% to 56.4%, demonstrating the effectiveness of RL-based post-training for agents.

Related event: Microsoft Releases Agent Lightning v1.0 for Reproducible Agent RL(2 posts)→

Original post →

More from coding & agent

coding & agent channel →