First runtime self-evolving agent achieves SOTA on SWE-bench
tom_doerr · x · 2026-08-15
Live-SWE-agent claims to be the first live, runtime self-evolving software engineering agent that expands and revises its own capabilities while working on issues.
Key Points:
- SOTA Performance: Achieves 79.2% on SWE-bench Verified with Claude Opus 4.5, leading current open-source scaffolds; scores 77.4% with Gemini 3 Pro.
- SWE-Bench Pro: Achieves a new state-of-the-art solve rate of 45.8%.
- Insight: Agents are software systems; modern LLM agents possess the intrinsic capability to extend or modify their own behavior at runtime.
More from coding & agent
- Test of 7 Coding Agents: Only Grok Accurately Assesses Vulnerability Severity — csuwildcat · 2026-08-15
- Building Cursor with Cursor: Where AI Fails and Human Review Is Needed — AI_Andrew · 2026-08-15
- Multi-Region Agent Failover Framework for Production Resilience — blaizedsouza · 2026-08-15
- Using Agent Trajectories to Mine Data for Model Distillation and Eval — hwchase17 · 2026-08-15
- Developer uninstalls AI agent Hermes after complex setup process — tristanbob · 2026-08-15
- Dev Accidentally Burns $1,475 in API Credits via Fast Mode Bug — bytebot · 2026-08-15