A Wild Async RL Config: 30 Steps x 25K Rollouts Per Step at Parallelism 4

willcb · x · 2026-09-22

A practitioner shared a notably aggressive RL training configuration: 30 training steps, 25,000 rollouts per step, with async parallelism of 4, calling it "a WILD config."

Configs at this rollout scale are typical of agent-style RL runs that need massive environment interaction, highlighting how async sampling infrastructure is being pushed for RL training.

Original post →

More from Infra

Infra channel →