From no benchmarks in the 1990s to no papers without them: a short history of ML evaluation
abursuc · x · 2026-09-16
A clip from the #ssad2026 talk: Kashyap Chitta traces the history of ML evaluation — in the 1990s almost no papers had benchmarks, while in the 2020s no paper ships without one. Establishing a common test framework culture was hard, but the resulting progress proved impressive.
More from Research
- Microsoft paper: unaligned small models 'launder' capabilities via frontier model consultation — dair_ai · 2026-09-16
- ModularRSI: Modular, benchmark-disjoint framework for generalizable agent harness self-improvement — IQuestLab · 2026-09-16
- 31,430-Trial Study: One Prompt Makes 11 Models Across OpenAI, Anthropic, Google Return Zero Bytes — rayanpal_ · 2026-09-16
- Robotics Team to Open-Source GALATEA, In-Hand Assembly and 6D Pose Control — ChongZzZhang · 2026-09-16
- CARLA veteran Ros shares synthetic data workflows to accelerate AV development — abursuc · 2026-09-16
- Grade AI like coworkers: open-source FrontierAgent framework ships with CLI TUI and fully local execution — aakashgupta · 2026-09-16