Default Codex CLI with GPT-5.5 scores 92.3% on XBOW, but the paper says the harness matters
evilsocket · x · 2026-07-24
- The post highlights a paper on autonomous penetration testing that evaluates coding agents on the 104-task XBOW benchmark.
- The authors run default Codex CLI, OpenCode, and Pi with the same GPT-5 model, same budget, same target interface, and same scoring rules to create a plain-agent baseline.
- The main takeaway is that specialized harnesses can inflate benchmark scores: the field should report model-matched baselines before crediting gains to architecture or scaffolding.
- The screenshot also shows the paper’s framing: controlled comparisons across plain-agent setups, prompt effects, and architecture residuals.
More from coding & agent
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- Steal this idea: prompt-to-hardware where agents assemble custom devices — paraschopra · 2026-09-11
- Model Is the Least Interesting Part: A Guide to Six Core AI Architectures from RAG to Multi-Agent — goyalshaliniuk · 2026-09-11
- Non-coder builds layered memory architecture: 20k tokens tracks a year of agent conversations — matteoianni · 2026-09-11
- Warp's six non-engineering teams all run on Linear and Claude Code — mon__lim · 2026-09-11