Scaling to 128 agents lifts pass rate from 19.3% to 28.8% on hardest ProgramBench tasks

OfirPress · x · 2026-09-29

The Agensh paper scales from 1 to 128 agents on the 5 hardest ProgramBench tasks, using GPT-5.6-sol (high reasoning) with Copilot as the single-agent harness. Mean 6-hour final pass rate rises from 19.31% to 28.78% — a 49% relative gain. Ofir Press highlights open questions: hierarchical vs flat organization, self-organized vs prescribed structures, and whether to pre-assign agent roles.

Related event: Microsoft Unveils Agensh, a Decentralized Multi-Agent Framework(2 posts)→

Original post →

More from coding & agent

coding & agent channel →