Microsoft proposes Harnessed Agentic RL to bridge training-deployment gap

SergioPaniego · x · 2026-08-20

Microsoft Research released Agent Lightning v1.0 and a paper proposing the "Harnessed Agentic RL" paradigm. This approach advocates training agents directly within their real-world deployment harnesses (which manage tools and context) rather than in reimplemented environments, thereby narrowing the gap between training and actual use. The paper details challenges unique to this method, such as retokenization, sample merging, advantage calculation, loss normalization, and backend scheduling, offering corresponding solutions.

Related event: Microsoft Releases Agent Lightning v1.0, Boosting SWE-bench to 56.4%(3 posts)→

Original post →

More from coding & agent

coding & agent channel →