New RL trick trades variance for bias, stabilizing tiny-batch training runs
willcb · x · 2026-09-27
The author introduces a new RL technique that trades variance for bias, offering higher stability compared to bias-free approaches in small runs with tiny batch sizes, which kept crashing in their experiments. They caveat that further scaling analysis is needed.
Related event: New RL trick trades bias for variance to stabilize small-batch training(2 posts)→
More from Research
- Ecological rationality paper to improve LLM cognitive bias benchmarks at NeurIPS — sethlazar · 2026-09-27
- NBER paper: rising cognitive skill productivity explains US within-occupation wage inequality — soumitrashukla9 · 2026-09-27
- Quail project on building with agents: stand on battle-tested community work — charles_irl · 2026-09-27
- 17M-parameter model beats frontier LLMs 80% of the time after cheap synthetic-data finetuning — max_paperclips · 2026-09-27
- Why the AI Scaling Hypothesis May Never Be Falsified: A Duhem–Quine Argument — burny_tech · 2026-09-27
- NYU's Tal Linzen cites two papers arguing tool use breaks Bender & Koller's 'no meaning' case — tallinzen · 2026-09-27