ProgramBench Called an Excellent, Far-from-Saturated Eval for Multi-Agent Coding Systems

jyangballin · x · 2026-09-28

Researcher jyangballin presents Agensh, the latest in his series of multi-agent × ProgramBench investigations. He argues ProgramBench (from FactoryAI's droid35719 and team) is an excellent long-horizon eval for SWE-agents and multi-agent coding systems — and that it is far from saturated, meaning substantial headroom remains for coding agents on long-horizon tasks.

Original post →

More from coding & agent

coding & agent channel →