DeLM project page: +17.5 points over strongest baseline with shared-context agents

rohanpaul_ai · x · 2026-10-08

The DeLM project page by Stanford researchers (Yuzhen Mao, Jerry Gu, et al.) details the mechanism: a shared context and task queue let agents coordinate asynchronously with no main agent. On Terminal-Bench 4.0, DeepSWE v1.1, and SWE-bench Verified, DeLM is faster and more accurate than Codex, Claude Code, their native subagents, and AOrchestra — up to 17.5 points more accurate than the strongest baseline and up to 2.49× faster than the harness it builds on. Back-end models: GPT-6-Astra (Codex) and Claude Opus 5.5 (Claude Code) at xhigh reasoning effort. Paper, code, traces, and plugin links included.

Related event: Stanford's DeLM Decentralized Multi-Agent Coding Runs 2.49x Faster(3 posts)→

Original post →

More from coding & agent

coding & agent channel →