TRL v1.15 ships fused LM head: 82% less peak VRAM, 7x longer sequences
QGallouedec · x · 2026-10-09
Hugging Face's TRL v1.15 is out, billed as its biggest optimization ever. The fused LM head, on by default, cuts peak VRAM by up to 82% and supports 7x longer sequences: DPO goes from 10k to 59k tokens, GRPO from 29k to 115k.
More from Infra
- GPUs already within 2x of brain efficiency, and still beat human workers on energy — MikePFrank · 2026-10-09
- Do AI agents still need Kubernetes? Berlin event says yes, with agent-on-K8s cases — Al_Grigor · 2026-10-09
- Cloud Backlogs Hit $1.69T, CoreWeave Posts $2.58B Quarter as Inference Becomes the Battleground — FinanceYF5 · 2026-10-09
- NVIDIA Is AI's Central Bank: A100 Paper Citations Still Beat H100+H200 Combined — FinanceYF5 · 2026-10-09
- Browser-Based Calculator Crunches the Real Power Cost of Self-Hosted LLMs vs Cloud APIs — paq85 · 2026-10-09
- PartyKit shuts down free hosted platform 2.5 years after Cloudflare acquisition — threepointone · 2026-10-09