RL config with 25k rollouts per step over just 30 steps called 'WILD'
andrew_n_carr · x · 2026-09-22
Andrew Carr flagged an unusual RL training setup: 30 steps with 25,000 rollouts per step at async-4, calling it a "WILD config" — a striking departure from typical RL training hyperparameters.
Related event: Wild RL training config sparks debate: 30 steps with 25k rollouts each(5 posts)→
More from Research
- Why OpenAI bets on math: it's the most verifiable domain for reinforcement learning — burny_tech · 2026-09-22
- Must-read papers of the week: recursive self-improvement, world models, KV cache compression — TheTuringPost · 2026-09-22
- Meta's A-MLE agent automates ML experimentation for ads ranking, cutting error 2.56% — rohanpaul_ai · 2026-09-22
- Blind RSA apps like Privacy Pass face real-world threat model from scaled oracle queries — matthew_d_green · 2026-09-22
- Harvard/MIT paper FINSKILLOPS makes financial AI self-improve via regression-tested skills — rohanpaul_ai · 2026-09-22
- Capping submissions per author won't cut much: volume comes from many low-output authors — furongh · 2026-09-22