Ruff author defends benchmarks: flawed but still the fairest way to compare AI models

EricBuess · x · 2026-09-27

Charlie Marsh (creator of Ruff) responds to criticism that the industry pivoted from hype-benching OpenAI scores to "benchmarks suck": today's benchmarks are fair to use but limited — he cites Anthropic's own note that the Opus vs Fable gap on paper exceeded real-world feel. Better benchmarks help everyone build and measure models, but it's extremely hard; models still competing on SWE-bench Verified alone would be bad for everyone.

Related event: Astral Founder Defends Benchmarks Amid "Useless" Backlash(2 posts)→

Original post →

More from Models

Models channel →