Astral founder defends benchmarks as useful but limited, admits better ones are hard
charliermarsh · x · 2026-09-27
Astral founder Charlie Marsh responded to a critic who called it "a very bearish sign" that after OpenAI spent weeks touting benchmark scores, the narrative pivoted to "benchmarks suck" — asking "were you wrong yesterday, or today?"\n\nMarsh said he didn't mean to dismiss benchmarks: evaluating models on today's benchmarks is fair, though they're limited. He noted Anthropic's Opus release, IIRC, acknowledged that the Fable–Opus gap in benchmarks felt wider than in real usage. He hopes the community builds better benchmarks — it helps everyone build better models and understand capabilities — but that turns out to be extremely hard.
Related event: Astral Founder Defends Benchmarks Amid "Useless" Backlash(2 posts)→
More from Models
- DeepSeek accused of benchmark gaming: new Flash model underwhelms in real use — Arindam_1729 · 2026-09-27
- GPT Agents Are Writing Eerie Uncategorizable Short Stories, and This One Is a Gem — RileyRalmuto · 2026-09-27
- Google ships a UI in AI Studio for trying its new audio and voice models — ammaar · 2026-09-27
- Rumors denied: Opus 4.5 not nerfed, renders 26,500-tile Byzantine unicorn mosaic — ctjlewis · 2026-09-27
- Opus 5.5 builds a fully procedural three.js world in one pass for ~$60 — EricBuess · 2026-09-27
- Devs say Claude Max plan offers 10-30x the value of the $200 Codex plan — chongdashu · 2026-09-27