Default Codex CLI with GPT-5.5 scores 92.3% on XBOW, sparking benchmark fatigue
moyix · x · 2026-07-24
A reply thread argues that the XBOW benchmark is already outdated, after a claim that a default Codex CLI setup with GPT-5.5 scored 92.3% on it.
- The poster says they did not build a custom harness or scaffolding, just used the default Codex CLI with GPT-5.5.
- The point is that continuing to advertise XBOW scores is misleading because the benchmark no longer reflects real performance.
- A follow-up note says the benchmark they published more than a year ago is now outdated and should no longer be used to measure performance.
More from coding & agent
- Google’s one-hour agentic engineering class covers memory, MCP, and multi-agent systems — HeyAmit_ · 2026-07-24
- Graph engineering gives Claude Code a more scalable workflow — blaizedsouza · 2026-07-24
- A Codex workflow uses the browser to run Deep Research across ChatGPT, Claude, and Gemini — JeremyNguyenPhD · 2026-07-24
- Grok Build 0.2.111 adds one-hour shell commands and more reliable MCP tools — mark_k · 2026-07-24
- AI agents need five layers, not just prompt engineering — blaizedsouza · 2026-07-24
- Builder rebuilds a Claude Code software factory around planning, implementation, review — blaizedsouza · 2026-07-24