TRL async GRPO adds LoRA sync via storage bucket and proxy, cutting 500-step training to 53 min
SergioPaniego · x · 2026-09-22
Hugging Face engineers detail LoRA support in TRL's AsyncGRPOTrainer (v1.14): only a few-MB adapter syncs to vLLM via a shared Storage Bucket instead of NCCL. Trainer and two vLLM replicas run as separate HF Jobs; a proxy routes each rollout to the replica holding its KV prefix. Iterating on AsyncGRPO metrics cut a 500-step run from 3h27m to 53m. Rationale: rank-1 LoRA suffices for policy-gradient RL per Thinking Machines' findings.
More from Infra
- bitsandbytes2 slightly delayed: zero-config lazy compression plus dynamic expert swaps for near-infinite KV cache — Tim_Dettmers · 2026-09-22
- Tim Dettmers releases runtime dynamic compression framework, hits 1.5-2.0 bit at high quality — Tim_Dettmers · 2026-09-22
- PSA: Non-US Users Should Consider Local AI in Case Governments Ban LLMs — TheMoonMidas · 2026-09-22
- Engram: A Local Encrypted Memory Vault Unifying Agent Memory Across AI Tools — Acceptable_Leg3950 · 2026-09-22
- DigitalOcean Managed Agents enters public preview with idle-pause billing — damianplayer · 2026-09-22
- Subconscious raises $5.1M to build an inference platform for long-horizon agents — CShorten30 · 2026-09-22