Inside ADS testing across 9 companies: sim-to-real gaps and no shared benchmarks

In the Driver's Seat: A Multi-Company Study on the Reality of Autonomous Driving System Testing

Qunying Song, Yuan Gao, Johannes Betz, Dietmar Pfahl, Mohammad Reza Mousavi, Federica Sarro

cs.SE

2026-07-17

Across 9 ADS companies in 6 countries: scenario and X-in-the-loop testing dominate, sim-to-real fidelity and scenario realism are the top bottlenecks, and shared benchmarks barely exist.

What problem this solves

Autonomous driving systems are shipping fast, but the industry has no agreed standard for how to test them or what "good enough" means. Which scenarios to run, which metrics to watch, which acceptance line to clear: each company decides for itself. Existing research tends to slice one face of the problem, a single technique or a single challenge, with no practitioner-level overview.

This is an interview study. The authors talked to 9 ADS testing practitioners at 9 companies across 6 countries (analysts, chief engineers, testers, product engineers, R&D leads), 45–75 minutes each, semi-structured, then thematic-coded the transcripts. Seven of the nine companies are automakers with over 10,000 employees, spanning China, Sweden, Germany, the UK, Japan, and Belgium. The output is a map of current practice, shared pain points, solution directions, and an evidence-centered closed-loop testing framework.

How they test

The dominant paradigm across the interviewed companies is scenario-based testing plus X-in-the-loop, swapping different components in and out of simulation at different stages:

On metrics, scenario coverage is named most often as the most important, followed by collision-related metrics (especially for parking), comfort (smooth steering, gentle braking), success rate (one company requires about 95% parking-space recognition before release), and fault rate. Acceptance lines are concrete: one requires over 90% accuracy across 300 km of open-road testing before sign-off.

What they found

There is no benchmark table, but the concentration of complaints is itself a result. Several pain points came up repeatedly:

One expert nailed the mileage metric: not all miles are equally valuable, and looping the same closed route is nothing like meeting diverse real situations. Another pointed at an open secret: suppliers market systems as L2/L2+/L2++ even when the experience is close to L3 or L4.

Solutions and the framework

The solution directions cluster around AI: 3D Gaussian Splatting for real-time scene reconstruction to raise simulation quality; end-to-end architectures that remove the explicit interface between perception and planning; world models that simulate sensor data directly; large models generating scenarios from real data; cloud-based continuous improvement to solve data logistics.

The most useful output is the evidence-centered closed-loop testing framework, in six stages: define safety claims and testing intent (using goal-structuring notation like GSN), build and maintain a scenario portfolio, plan evidence production and route to the right test environment (MIL/SIL/HIL/VIL), apply evidence gates and acceptance criteria, build the safety argument and decide progression or release, then close the loop by learning from operation and feeding data back into testing. The core claim is that testing cannot be just mileage accumulation; it has to be a traceable chain of evidence, each step accumulating verifiable proof for "the system is safe under these conditions."

Why it matters

For anyone in autonomous-driving perception, planning, or simulation, this study maps what the industry actually worries about. Sim-to-real, scenario realism, world models, and end-to-end are hot in academia right now, and they happen to line up exactly with the sorest points on the testing floor. If your work shrinks the sim-to-real gap or makes generated scenarios more credible, there is real industrial demand waiting.

Limitations

This is a qualitative study of 9 people at 9 companies, a small sample, and visibly skewed toward European and Chinese automakers. The US players (Waymo, Cruise and the like) are absent, yet those are precisely the companies furthest along on public benchmarks and mileage disclosure. The authors contacted 21 companies and 9 participated, so self-selection bias is real: the ones willing to talk may already care more about process. Interviews reflect only what respondents said, with no on-site verification of their actual workflows. The framework itself is synthesized from interviews and has not been validated in a real project to show it actually improves testing efficiency or safety.

Terms

Source

Related papers

All paper explainers