Dev critique: step-like gains on a driving benchmark are misleading
ben_j_todd · x · 2026-09-24
Ben Todd highlights a useful comment from Substack (by Aman Karunakaran) arguing that the "step-like capability jumps" shown on a remote-driving-style benchmark are misleading.
The benchmark gives a model 3 attempts and measures how far along a cone-marked course it can go before going off track. But the course's difficulty is not monotonically increasing — it contains only a few chokepoints. Once a model clears a chokepoint, it usually progresses until the next one. Comparing Fable (orange) vs Astra (purple), the commenter notes Fable's third attempt and Astra's first attempt look essentially the same; in the video, some cones are sparse and hard to see later in the course, and the course opens with a very sharp turn that most models fail, including Fable's first two attempts.
More from Models
- GPT-6 Sol and Luna Score Below GPT-5.6: First Major Release Weaker Than Its Predecessor — srchvrs · 2026-09-24
- OpenAI launches GPT-6 Sol and Luna with 50% lower API prices than GPT-5.6 — DeryaTR_ · 2026-09-24
- NVIDIA open-sources Nemotron 3 Diarization, #1 on VoiceArena with 14.72% DER — solyarisoftware · 2026-09-24
- Opus 5.5 Is 10x Slower Than Fable for Financial Modeling, User Reports — JOBhakdi · 2026-09-24
- OpenAI Supercharges ChatGPT Voice: Plugins for Email, Calendar, Slack, Powered by GPT-6 — Dimillian · 2026-09-24
- Code Arena launches WebDev leaderboard for comparing coding models — arena · 2026-09-24