FlashREINFORCE: open-source single-rollout async LLM RL with 6,000+ stable updates
JFPuget · x · 2026-09-14
FlashREINFORCE claims to be the first open-source critic-free, single-rollout asynchronous LLM RL method with 6,000+ stable updates, combining One-Batch REINFORCE, Sequence Trust Region, and sample-mean optimization.
@giffmana breaks down the mechanics: unlike GRPO's 32×4 grouping, each prompt yields one rollout, 128 rollouts form a batch used for a single optimizer step.
The key win is handling stragglers on long rollouts: with no groups, no rollout waits on siblings — 500 runners can feed a queue and the trainer just batches the next 128 that arrive, regardless of order.
Paper and code are open-sourced.
Related event: NVIDIA Open-Sources FlashREINFORCE, a Critic-Free Async RL Framework(3 posts)→
More from Research
- New Data Agent Benchmark: Best Frontier Agent Passes Just a Third of 54 Multi-DB Queries — CShorten30 · 2026-09-14
- ActionSplice edits actions in video world models without replaying evaluations — TexasAMUniversity · 2026-09-14
- Kamoun Lab Publishes 10 Lessons on AI Protein Structure Prediction — AllThingsApx · 2026-09-14
- 1.2M STEM dissertations show government is the top funder of frontier PhD research — joshgans · 2026-09-14
- NBER Economics of AI 2026 heads to Toronto with focus on AI in China — joshgans · 2026-09-14
- Fruit fly brain trained 35 min on one H100 drives real traffic with zero collisions, open-sourced — TinfoilTricorn · 2026-09-14