"Everyone is cheating on AI benchmarks": a hacker's plea to optimize for the real world
hackgoofer · x · 2026-09-14
After a hackathon, hackgoofer delivered a spicy take: stop optimizing for benchmarks and make things actually work in the real world. Citing an article titled "Lies, Damned Lies, and Benchmarks," he argues intelligence is jagged partly because everyone is gaming benchmarks—since they're the easiest thing to auto-improve on. Refusing that easy path takes guts, insight and time. His team, with a release coming soon, is making precommitments to do better.
More from Models
- Users say mystery model 'instinct' outperforms Grok and Muse in hands-on tests — Scobleizer · 2026-09-14
- Kokoro TTS ported to Apple Core AI: 54 voices running fully on-device with zero API cost — amos_gyamfi · 2026-09-14
- Multi-level simulation evals may be uninformative after Anthropic's Hacker Opus post — herbiebradley · 2026-09-14
- Dev Says 'Astra' Via OpenRouter Burned $30 in About 9 Seconds — haydendevs · 2026-09-14
- Dev Shares Token Split: Open Models Outused Closed 10:1 for Coding and Research — xeophon · 2026-09-14
- Dario Amodei: Claude has spotted medical issues doctors missed, big gains in biology — Olivier__OG · 2026-09-14