ProgramBench Called an Excellent, Far-from-Saturated Eval for Multi-Agent Coding Systems
jyangballin · x · 2026-09-28
Researcher jyangballin presents Agensh, the latest in his series of multi-agent × ProgramBench investigations. He argues ProgramBench (from FactoryAI's droid35719 and team) is an excellent long-horizon eval for SWE-agents and multi-agent coding systems — and that it is far from saturated, meaning substantial headroom remains for coding agents on long-horizon tasks.
More from coding & agent
- Claude Code creator Boris Cherny: bet on general models, skip fine-tuning — rohanpaul_ai · 2026-09-28
- Personal agents need fixed chores, not more tokens: define boundaries before letting them run — sujingshen · 2026-09-28
- App devs face two paths in the personal-agent era: integrate MCP or become the agent — sujingshen · 2026-09-28
- 30 verified Claude Opus 5.5 browser animation cases, ranked by views, with prompts — dotey · 2026-09-28
- One agent per household: whose veto wins when family calendars conflict? — sujingshen · 2026-09-28
- Rolldown to ship experimental inlineCommonChunks to cut small shared chunks — cnakazawa · 2026-09-28