Benchmarks let people run with preferred narratives — capability vectors beat leaderboards
GlenBradley · x · 2026-09-04
Glen Bradley argues benchmark reporting around the new Astra model has gotten out of hand: half the numbers make it look like an average frontier model, half make it look far ahead — people pick benchmarks that support the story they already want to tell.
He leans optimistic but admits a single frontier-model leaderboard is increasingly meaningless. His proposal: describe models as capability vectors — {reasoning, coding, research breadth, research completeness, long-horizon coherence, tool use, factuality, multimodal reasoning, latency, cost…} — so comparisons become surfaces rather than rankings. A 5% gain on a short multiple-choice test and a 60% gain on a large evidence-backed investigation are not economically comparable.
More from Models
- Martian says routing across 44 LLMs cuts errors 46% vs best single model on 16 benchmarks — Arindam_1729 · 2026-09-04
- Ex-OpenAI safety lead Miles Brundage: if your primary emotion on AI isn't concern, you're misreading it — Miles_Brundage · 2026-09-04
- Gary Marcus on GPT-6 Astra: symbolic world models are vindication, but no proof of AGI — GaryMarcus · 2026-09-04
- Ex-Google X Quant Guillaume Verdon Says OpenAI Is 'Kinda Back' — But Vibes, Not Benchmarks, Will Decide — beffjezos · 2026-09-04
- Mystery model "Astra" reportedly beats 5.6 Sol Pro on FrontierMath T4 — ctjlewis · 2026-09-04
- Inkling and Inkling-Small models get a free inference tier, and it's staying — simonguozirui · 2026-09-04