Benchmark Score Jumps Misleading: Critique of Cone-Course Model Evaluation
ben_j_todd · x · 2026-09-24
Ben Todd shares a Substack critique of a model's cone-course benchmark "step-like" score increase.
- Scoring gives the model 3 attempts and measures distance along the cone course; difficulty doesn't scale monotonically — progress hinges on a few chokepoints rather than gradual skill.
- Fable's third attempt nearly matches Astra's first; some cones are sparse/hard to see, and a sharp opening turn trips most models, including Fable's first two runs.
- Todd adds: the result isn't great yet, but at this rate of progress it could be within a year.
Related event: Google Astra Appears to Drive Without Dedicated Training(4 posts)→
More from Models
- Artificial Analysis launches TTS leaderboard with new Pronunciation Robustness Benchmark across 95 models — ArtificialAnlys · 2026-09-24
- Most users treat stochastic AI as deterministic, ignoring false positives — tokenbender · 2026-09-24
- Human text flagged 100% AI: why people treat stochastic detectors as oracles — tokenbender · 2026-09-24
- GPT-6 Sol and Luna Score Below GPT-5.6: First Major Release Weaker Than Its Predecessor — srchvrs · 2026-09-24
- Reddit users slam GPT-6 Sol and Luna as dumber than 5.6 despite cheaper pricing — Obvious_Resort8887 · 2026-09-24
- Opus 5.5 Is 10x Slower Than Fable for Financial Modeling, User Reports — JOBhakdi · 2026-09-24