Default Codex CLI with GPT-5.5 scores 92.3% on XBOW, but the paper says the harness matters
evilsocket · x · 2026-07-24
- The post highlights a paper on autonomous penetration testing that evaluates coding agents on the 104-task XBOW benchmark.
- The authors run default Codex CLI, OpenCode, and Pi with the same GPT-5 model, same budget, same target interface, and same scoring rules to create a plain-agent baseline.
- The main takeaway is that specialized harnesses can inflate benchmark scores: the field should report model-matched baselines before crediting gains to architecture or scaffolding.
- The screenshot also shows the paper’s framing: controlled comparisons across plain-agent setups, prompt effects, and architecture residuals.
More from coding & agent
- Claude Code vs Codex: the harness may matter more than the model — al_kRicha · 2026-07-24
- Which model should power each agent task in production? — Logical-Silver-272 · 2026-07-24
- Offloop says its 4-person multi-agent harness beat Claude Code and Codex on GDPval — rohanpaul_ai · 2026-07-24
- Graph engineering, in one line: deterministic multi-agent loops become DAGs — vincent_koc · 2026-07-24
- BackSearch’s development was led by @ranko3000 with feedback from several collaborators — rosstaylor90 · 2026-07-24
- BackSearch lets agents query the web as-of a past date to avoid future leakage — brianrkelly · 2026-07-24