Customer calls out OpenAI: benchmarks were great last week, now "benchmarks suck"?
macintogdev · x · 2026-09-27
A developer directly challenged the narrative on X: right after OpenAI spent weeks crowing about benchmark results (which "don't come close to matching user experiences"), voices are now pivoting to "benchmarks suck" — being read as a very bearish sign. He stresses he's a customer, not a hater, and genuinely asks: were you all wrong yesterday, or today?\n\nThe post drew a reply from Astral founder Charlie Marsh, who defended benchmarks as useful but limited and called for building better ones.
More from Models
- DeepSeek accused of benchmark gaming: new Flash model underwhelms in real use — Arindam_1729 · 2026-09-27
- GPT Agents Are Writing Eerie Uncategorizable Short Stories, and This One Is a Gem — RileyRalmuto · 2026-09-27
- Google ships a UI in AI Studio for trying its new audio and voice models — ammaar · 2026-09-27
- Rumors denied: Opus 4.5 not nerfed, renders 26,500-tile Byzantine unicorn mosaic — ctjlewis · 2026-09-27
- Opus 5.5 builds a fully procedural three.js world in one pass for ~$60 — EricBuess · 2026-09-27
- Devs say Claude Max plan offers 10-30x the value of the $200 Codex plan — chongdashu · 2026-09-27