E2E driving evaluation shifts: nuPlan open-loop flaws exposed, NAVSIM benchmarks drove two years of progress
abursuc · x · 2026-09-16
- At the #ssad2026 workshop, researchers reviewed the evolution of end-to-end autonomous driving evaluation: nuPlan's open-loop protocol has known flaws, and the paper "Is Ego Status All You Need for Open-Loop Autonomous Driving" exposed how ego status can shortcut the benchmark.
- NAVSIM then emerged as a lightweight-simulation benchmark suite, acting as a super-enabler for e2e driving model progress over the last two years, though not without imperfections.
- Broader observation: in the 1990s no papers had benchmarks; in the 2020s no papers ship without them.
More from Research
- HarnessVLN: training-free embodied navigation agent sets SOTA on four benchmarks — Yang Chen · 2026-09-16
- TROT: Tsallis-Regularized Optimal Transport Unifies Wasserstein and KL Divergences — FrnkNlsn · 2026-09-16
- Professor estimates viral post-training algorithms work out of the box only ~5% of the time — Kangwook_Lee · 2026-09-16
- AlpaSim Challenge Borrows LLM Multi-Domain Benchmarking, Uses Item Response Theory for Autonomous Driving Evals — abursuc · 2026-09-16
- CoLLAs 2026 Keynote: Continual Model Merging via Subspace Modeling and Low-Rank Experts — apsarathchandar · 2026-09-16
- Interpretability researcher lists top open problems in decoding model activations — wesg52 · 2026-09-16