Distributed RL Post-Training Achieved on 14 Macs

erfan_mhi · reddit · 2026-07-16

Pluralis Research shared a distributed RL post-training experiment where 14 consumer-grade Macs across 4 countries handled rollout generation, while a remote B200 managed bf16 gradient updates.

Key Approaches

Results

They evaluated Stoa on PaperSearchQA, a multi-turn biomedical retrieval task:

Conclusion & Next Steps

The author emphasizes that consumer hardware can shoulder the bulk of rollout costs. Combining this Stoa approach with Agoro—which previously performed pre-training on hundreds of consumer GPUs—could enable migrating larger-scale inference and training entirely onto user-owned hardware.

Related event: Pluralis Runs Distributed RL Post-Training on 14 Consumer Macs(2 posts)→

Original post →

More from Infra

Infra channel →