Early hands-on: GPT-6-Sol xhigh underwhelms on Codex's Boeing bench
victormustar · x · 2026-09-23
@victormustar reports that running GPT-6-Sol xhigh in Codex on the "Boeing bench" yields unimpressive results. No numbers are given in the text itself, with details in the attached screenshot. An early negative datapoint on the newest model's coding ability.
More from Models
- Dev claims Opus 5.5 generated stunning animation frames with pure code, no game engine — RileyRalmuto · 2026-09-23
- Frontier releases suggest open models like Kimi K3 are severely undertrained — zeeshanp_ · 2026-09-23
- MiMo V2.6 Pro and Flash take top two spots on Vals Index open-weight leaderboard — burny_tech · 2026-09-23
- ChatGPT 3.5 felt as good as 5.6 in memory — like replaying a childhood game — flowersslop · 2026-09-23
- Opus 5.5 vs Sol 6 one-shot coding test: "it is not close" — alvelda · 2026-09-23
- TestingCatalog adds email AI brief as GPT-6 Sol/Luna and Opus 5.5 land — testingcatalog · 2026-09-23