TRL Adds Async GRPO with LoRA Weight Sync over HF Buckets, Cutting Training from 3.5h to 53min
_lewtun · x · 2026-09-18
Hugging Face engineer Amine Dirhoussi announced LoRA training in TRL's AsyncGRPOTrainer, with RL weight sync over HF Buckets.
Key details:
- Trainer and vLLM inference run on separate HF Jobs machines — no NCCL, no shared disk.
- Only the small LoRA adapter delta (few MB for a 1.5B model) is synced, mounted at the same path in every Job.
- The same reward now takes 53 minutes instead of 3h27.
The setup lets small teams run async GRPO entirely on HF infrastructure.
More from Infra
- Starlink connects thousands of rural Latin American schools, from 10,000 antennas in Honduras to Bolivia's national rollout — XFreeze · 2026-09-18
- 0.05% sampling to validate cache hits: developer marvels at compute saved across the system — DanielLockyer · 2026-09-18
- Payments firms race to own AI inference: Stripe taps OpenRouter, Ramp enters the chain — xkonjin · 2026-09-18
- OpenDCAI/DataFlow: open-source pipeline toolkit for pre-training data prep — Puzzleheaded_Box2842 · 2026-09-18
- Qdrant wraps 4+ hour Vector Space Stream on vector search — recording now live — qdrant_engine · 2026-09-18
- Investor: QNX's microkernel is an undervalued moat for the agentic AI era — pdamodaran · 2026-09-18