Mean scores mislead RL progress; Google Research proposes IQM and confidence intervals
burkov · x · 2026-08-26
Deep reinforcement learning experiments are expensive, leading researchers to compare algorithms using few runs despite high inter-run variance. A paper from Google Research and Université de Montréal uses large-scale Atari experiments to demonstrate how ordinary mean or median scores give a misleading picture of progress. The authors propose practical reporting methods for budget-constrained scenarios: confidence intervals, performance profiles, and the interquartile mean (IQM).
More from Research
- Scaling Laws Extrapolation Unreliable Due to Statistical Reasons — DanielKhashabi · 2026-08-26
- KostasVisualizations: Interactive Visualizations for Computer Vision and Machine Learning — CSProfKGD · 2026-08-26
- AI Scholar Praises Model Architecture Research: Value of 180B Scale Validation and Ablation Studies — kastnerkyle · 2026-08-26
- FixAnything: Unified 3D Rendering Artifact Repair via Video Generative Priors — gabriel1 · 2026-08-26
- Paper proposes a fix for running context recirculation issues — yogthos · 2026-08-26
- Study Finds Attention-Mask Checks Miss Causality Leaks in Sequence Models — VIDraft · 2026-08-26