DeLM's decentralized multi-agent system runs 2.49x faster, but MAS evals pick wildly different metrics

jyangballin · x · 2026-10-08

Commenting on the updated DeLM paper, the author highlights a decentralized multi-agent system where parallel agents coordinate via shared context and a task queue with no main agent: agents share discoveries, reuse intermediate results, and correct each other. Evaluation now spans SWE-bench, Terminal-Bench 4.0, DeepSWE v1.1, and ProgramBench.

The notable takeaway is how different MAS proposals on ProgramBench showcase success on different axes:

The author draws parallels to the early SWE-bench harness era (Devin, SWE-agent, Agentless, AutoCodeRover, OpenHands), when each team picked a different efficiency axis to claim superiority — and notes that adding more agents often doesn't actually help solve tasks more efficiently.

Original post →

More from coding & agent

coding & agent channel →