TRL Adds Async GRPO with LoRA Weight Sync over HF Buckets, Cutting Training from 3.5h to 53min

_lewtun · x · 2026-09-18

Hugging Face engineer Amine Dirhoussi announced LoRA training in TRL's AsyncGRPOTrainer, with RL weight sync over HF Buckets.

Key details:

The setup lets small teams run async GRPO entirely on HF infrastructure.

Original post →

More from Infra

Infra channel →