Scaling to 128 agents lifts pass rate from 19.3% to 28.8% on hardest ProgramBench tasks
OfirPress · x · 2026-09-29
The Agensh paper scales from 1 to 128 agents on the 5 hardest ProgramBench tasks, using GPT-5.6-sol (high reasoning) with Copilot as the single-agent harness. Mean 6-hour final pass rate rises from 19.31% to 28.78% — a 49% relative gain. Ofir Press highlights open questions: hierarchical vs flat organization, self-organized vs prescribed structures, and whether to pre-assign agent roles.
Related event: Microsoft Unveils Agensh, a Decentralized Multi-Agent Framework(2 posts)→
More from coding & agent
- Sonnet 5.5 clones open-source editor Proof at low effort, joining elite group of just four models — every · 2026-09-29
- Agent coding at scale hits a CI bottleneck: dev built dsr after $5k/month GitHub Actions bill — doodlestein · 2026-09-29
- Cloud Planner + Local Coder: 2.7x Cheaper but 9x Slower in a Real Agent Build — GapNew4766 · 2026-09-29
- Retail Investor Builds a Semi-Automated Stock Trading Script with Claude and a Graham Formula — Elfwood86 · 2026-09-29
- Turbopuffer cofounder confirms Cursor dropped server-side RAG for grep as models improved — ivan_bezdomny · 2026-09-29
- Dev builds Skyfall Rush with Claude Opus in 2 days, playable free in browser — TAbrodi · 2026-09-29