Default Codex CLI with GPT-5.5 scores 92.3% on XBOW, sparking benchmark fatigue

moyix · x · 2026-07-24

A reply thread argues that the XBOW benchmark is already outdated, after a claim that a default Codex CLI setup with GPT-5.5 scored 92.3% on it.

Original post →

More from coding & agent

coding & agent channel →