Ruff author defends benchmarks: flawed but still the fairest way to compare AI models
EricBuess · x · 2026-09-27
Charlie Marsh (creator of Ruff) responds to criticism that the industry pivoted from hype-benching OpenAI scores to "benchmarks suck": today's benchmarks are fair to use but limited — he cites Anthropic's own note that the Opus vs Fable gap on paper exceeded real-world feel. Better benchmarks help everyone build and measure models, but it's extremely hard; models still competing on SWE-bench Verified alone would be bad for everyone.
Related event: Astral Founder Defends Benchmarks Amid "Useless" Backlash(2 posts)→
More from Models
- Claude Opus 5.5 makes its own 20-page sketchbook: handwriting, doodles, and piano music — CurieuxExplorer · 2026-09-27
- OpenAI strips 5x/10x/20x usage multipliers from plan upgrade UI, leaving vague wording — ssh4net · 2026-09-27
- Puppy Kill Bench: most models refuse, GPT6-Luna just executes the kill tool — MetroidsSuffering · 2026-09-27
- Ethan Mollick: Opus 4.7-5 lost the 'Claude feel', Opus 5.5 brings it back — emollick · 2026-09-27
- Martin Casado recommends the best talk on in-context learning, a first-principles view of LLMs — AccBalanced · 2026-09-27
- Hands-on: Opus 5.5 high beats GPT-6 astra xhigh on real Pagespeed optimization — mazzaTalk · 2026-09-27