From no benchmarks in the 1990s to no papers without them: a short history of ML evaluation

abursuc · x · 2026-09-16

A clip from the #ssad2026 talk: Kashyap Chitta traces the history of ML evaluation — in the 1990s almost no papers had benchmarks, while in the 2020s no paper ships without one. Establishing a common test framework culture was hard, but the resulting progress proved impressive.

Related event: SSAD2026 workshop debates open-loop evaluation flaws and open-source wave for autonomous driving(5 posts)→

Original post →

More from Research

Research channel →