DeLM's decentralized multi-agent system runs 2.49x faster, but MAS evals pick wildly different metrics
jyangballin · x · 2026-10-08
Commenting on the updated DeLM paper, the author highlights a decentralized multi-agent system where parallel agents coordinate via shared context and a task queue with no main agent: agents share discoveries, reuse intermediate results, and correct each other. Evaluation now spans SWE-bench, Terminal-Bench 4.0, DeepSWE v1.1, and ProgramBench.
The notable takeaway is how different MAS proposals on ProgramBench showcase success on different axes:
- Opus model cards: accuracy vs. tokens
- DeLM: accuracy vs. wall-clock time (2.49x faster, but higher total cost)
- Agensh: accuracy vs. number of agents at fixed time
The author draws parallels to the early SWE-bench harness era (Devin, SWE-agent, Agentless, AutoCodeRover, OpenHands), when each team picked a different efficiency axis to claim superiority — and notes that adding more agents often doesn't actually help solve tasks more efficiently.
More from coding & agent
- Empirical eval shows AI-written unit and integration tests don't improve agent success rates — GabGarrett · 2026-10-08
- Agency Agents hits 158k GitHub stars: an entire AI company of role-specialized agents — sujingshen · 2026-10-08
- Musk promotes Grok Bot: X Corp.'s AI coworker app that logs in and works your tools 24/7 — elonmusk · 2026-10-08
- Demo is not delivery: what enterprise AI adoption actually takes — sujingshen · 2026-10-08
- Musk endorses Grok bot + Cursor cloud agents combo as a major productivity boost — elonmusk · 2026-10-08
- Dev: in the AI coding era, tabs-vs-spaces and style nitpicking never mattered — facontidavide · 2026-10-08