160,000 Training Runs Across 114 Datasets Show No Algorithm Dominates Offline Policy Learning

raivn · hf · 2026-10-08

A large-scale empirical study of offline reinforcement and imitation learning trains over 160,000 policies across 114 datasets.

Key findings:

They release JumpStart: a resource suite with all trained policies, per-model scores and hyperparameters, strong baselines, training/eval code, and an extensible website.

Original post →

More from Research

Research channel →