Astral founder defends benchmarks as useful but limited, admits better ones are hard

charliermarsh · x · 2026-09-27

Astral founder Charlie Marsh responded to a critic who called it "a very bearish sign" that after OpenAI spent weeks touting benchmark scores, the narrative pivoted to "benchmarks suck" — asking "were you wrong yesterday, or today?"\n\nMarsh said he didn't mean to dismiss benchmarks: evaluating models on today's benchmarks is fair, though they're limited. He noted Anthropic's Opus release, IIRC, acknowledged that the Fable–Opus gap in benchmarks felt wider than in real usage. He hopes the community builds better benchmarks — it helps everyone build better models and understand capabilities — but that turns out to be extremely hard.

Related event: Astral Founder Defends Benchmarks Amid "Useless" Backlash(2 posts)→

Original post →

More from Models

Models channel →