ProgramBench multi-agent eval: Opus 5.5 fastest with a 5-agent team, Sonnet 5.5 with subagents
jyangballin · x · 2026-09-29
In a multi-agent evaluation on ProgramBench, the fastest configurations differ by model: Sonnet 5.5 performed best with subagents, while Opus 5.5 hit top speed using a 5-agent team. A useful data point that optimal agent orchestration is model-specific.
More from coding & agent
- Devin cuts prices up to 70% while topping FrontierCode 1.1 Extended leaderboard — brandon_galang · 2026-09-29
- Doug Turnbull's Search Training Kicks Off Next Week, Focused on Agentic Retrieval and RAG — JnBrymn · 2026-09-29
- Live Walkthrough of DSPy's Jev Integration and ReAnchor Optimizer Set for Wednesday — dbreunig · 2026-09-29
- The AI 'architect' ignored all requests, charged 11 million quid, so he drew it himself — snikolov · 2026-09-29
- Dev burns all his Opus 5.5 tokens building a code-only three.js procedural animal pack — majidmanzarpour · 2026-09-29
- Multi-model coding: routing Opus, GPT, Kimi, DeepSeek in one Codex session — omarsar0 · 2026-09-29