Mean scores mislead RL progress; Google Research proposes IQM and confidence intervals

burkov · x · 2026-08-26

Deep reinforcement learning experiments are expensive, leading researchers to compare algorithms using few runs despite high inter-run variance. A paper from Google Research and Université de Montréal uses large-scale Atari experiments to demonstrate how ordinary mean or median scores give a misleading picture of progress. The authors propose practical reporting methods for budget-constrained scenarios: confidence intervals, performance profiles, and the interquartile mean (IQM).

Original post →

More from Research

Research channel →