Frontier-Model Bug Benchmark: 105 Problems Models Failed at the Start of 2026
PawelHuryn · x · 2026-09-23
Commenting on a frontier-model bug-solving benchmark, PawelHuryn clarifies: the 105 bugs aren't random but hard problems frontier models failed to solve at the start of 2026; unplanted bugs aren't counted since OpenAI 'bugmaxes' by reporting many real, theoretical, and irrelevant issues whose solutions would complicate results—arguably negative; judges from different model families score solutions against a secret answer key, and identified-but-unfixed issues are tracked.
More from Models
- GPT-6 Sol and Luna already usable in Codex, early user reports — airesearch12 · 2026-09-23
- GPT-6 Sol scores slightly below GPT-5.6 Sol on DeepSWE, only cheaper — Angaisb_ · 2026-09-23
- Matt Shumer on Opus 5.5: 'feels like a much smarter Opus 4.6' — mattshumer_ · 2026-09-23
- GPT-6 Sol claimed to cost 50% less than GPT-5.6 Sol — cedric_chee · 2026-09-23
- Four frontier models in days: Grok 4.7, Opus 5.5, GPT-6 Sol and Luna — msg · 2026-09-23
- OpenAI reportedly rolling out GPT-6 Sol and GPT-6 Luna on ChatGPT, Codex, and APIs — testingcatalog · 2026-09-23