160,000 Training Runs Across 114 Datasets Show No Algorithm Dominates Offline Policy Learning
raivn · hf · 2026-10-08
A large-scale empirical study of offline reinforcement and imitation learning trains over 160,000 policies across 114 datasets.
Key findings:
- No algorithm dominates: aggregate performance among the strongest methods is often close, but leaders differ substantially across environments
- Proper hyperparameter tuning frequently reshuffles perceived algorithm rankings
- Benchmark composition can produce conflicting conclusions
- The authors study hyperparameter sensitivity and transfer, identifying a simple strategy for deriving strong default configurations
- A dataset-conditioned recommender provides task-specific algorithm recommendations
They release JumpStart: a resource suite with all trained policies, per-model scores and hyperparameters, strong baselines, training/eval code, and an extensible website.
More from Research
- Bittensor's SN107 Lets Miners Earn by Running AI Agents to Produce Genomic Data — markjeffrey · 2026-10-08
- CoLM 2026 Poster: Vibe-Voting LLMs and Why Users Distrust Benchmarks — boknilev · 2026-10-08
- Baseten's Base Labs has all 3 papers accepted at NeurIPS workshops — baseten · 2026-10-08
- Open-vocabulary text search fused into Google 3D Tiles across 27M voxels of San Francisco — bilawalsidhu · 2026-10-08
- SEAR research project to debut at COLM 2026, paper and code coming soon — WenhuChen · 2026-10-08
- LeJEPA lets you pretrain self-supervised vision models on your own data — randall_balestr · 2026-10-08