Dev critique: step-like gains on a driving benchmark are misleading

ben_j_todd · x · 2026-09-24

Ben Todd highlights a useful comment from Substack (by Aman Karunakaran) arguing that the "step-like capability jumps" shown on a remote-driving-style benchmark are misleading.

The benchmark gives a model 3 attempts and measures how far along a cone-marked course it can go before going off track. But the course's difficulty is not monotonically increasing — it contains only a few chokepoints. Once a model clears a chokepoint, it usually progresses until the next one. Comparing Fable (orange) vs Astra (purple), the commenter notes Fable's third attempt and Astra's first attempt look essentially the same; in the video, some cones are sparse and hard to see later in the course, and the course opens with a very sharp turn that most models fail, including Fable's first two attempts.

Original post →

More from Models

Models channel →