The jagged frontier of LLM capability is just a map of who built an automated grader
jessi_cata · x · 2026-09-14
Sharing Jon Stokes' article 'Everyone should be extremely skeptical of AI benchmarks', jessicata highlights its core claim: the jagged frontier of LLM capability isn't an inherent property of intelligence but a map of which tasks somebody figured out how to write an automated grader for. The article argues benchmark narratives diverge systematically from real capability.
More from Models
- Codex 5x users report odd usage drain: ~1% every 10 minutes and 10-day reset cycles — Andreeez · 2026-09-14
- Claude turns out to be able to draft its own bug reports — Aizkmusic · 2026-09-14
- GPT-6 Astra: 1M-token context, 2.5x price, and a 'Critical' cyber risk rating — Thirumalaivasan_GJ · 2026-09-14
- Harvey Trains Own Legal Model on Kimi K3, Signaling a Shift to Enterprise Self-Training — 机器之心 · 2026-09-14
- GPT Image 2.5 is live in ChatGPT, SOTA on all image leaderboards — sherwinwu · 2026-09-14
- Dev Praises Qwen3.8-27B After Week of Testing: One-Shots Vague Prompts, Writes Own Tests — DustNearby2848 · 2026-09-14