New RL trick trades variance for bias, stabilizing tiny-batch training runs

willcb · x · 2026-09-27

The author introduces a new RL technique that trades variance for bias, offering higher stability compared to bias-free approaches in small runs with tiny batch sizes, which kept crashing in their experiments. They caveat that further scaling analysis is needed.

Related event: New RL trick trades bias for variance to stabilize small-batch training(2 posts)→

Original post →

More from Research

Research channel →