Microsoft's Agent Lightning boosts SWE-bench score to 56.4% with 6K samples
omarsar0 · x · 2026-08-19
Microsoft released Agent Lightning v1.0, a framework connecting agent 'harnesses' (which own tools, context, and control flow) to Reinforcement Learning training via an endpoint proxy. It addresses engineering challenges like retokenization, sample merging, and advantage calculation.
Using only 6K training examples and modest compute, the framework improved Qwen2.5-9B's performance on SWE-bench Verified from 41.8% to 56.4%, demonstrating the effectiveness of RL-based post-training for agents.
Related event: Microsoft Releases Agent Lightning v1.0 for Reproducible Agent RL(2 posts)→
More from coding & agent
- Using Codex and NeMo-RL to Train Models on Slurm Clusters — JFPuget · 2026-08-19
- Consumer agents got so good that email and calendar apps look disruptable — omooretweets · 2026-08-19
- Stripe integrates with Grok Bot for agentic customer support operations — jeff_weinstein · 2026-08-19
- Multi-Agent Workflow: Splitting Tasks Beats One全能 Agent for Dev — RonnySaya · 2026-08-19
- Upcoming Agent course series: First video on evaluating evals — ben_burtenshaw · 2026-08-19
- Claude Code plugin rewrites output to plain English using local LLMs — tom_doerr · 2026-08-19