Benchmarks let people run with preferred narratives — capability vectors beat leaderboards

GlenBradley · x · 2026-09-04

Glen Bradley argues benchmark reporting around the new Astra model has gotten out of hand: half the numbers make it look like an average frontier model, half make it look far ahead — people pick benchmarks that support the story they already want to tell.

He leans optimistic but admits a single frontier-model leaderboard is increasingly meaningless. His proposal: describe models as capability vectors — {reasoning, coding, research breadth, research completeness, long-horizon coherence, tool use, factuality, multimodal reasoning, latency, cost…} — so comparisons become surfaces rather than rankings. A 5% gain on a short multiple-choice test and a 60% gain on a large evidence-backed investigation are not economically comparable.

Original post →

More from Models

Models channel →