TRL v1.15 defaults to fused LM head, extending training sequences up to 6.9x

LysandreJik · x · 2026-10-09

Hugging Face shipped TRL v1.15, a memory-efficiency-focused release: SFT, DPO, KTO, GRPO, RLOO and distillation now default to a fused LM head — instead of materializing the huge [batch, seq, vocab] logits tensor, a Triton kernel computes token-level quantities directly.

Benchmarks (Gemma 3 1B, 262k vocab, same GPU):

Up to 6.9x longer sequences, peak memory at 8k context cut by 52-82%, and training is 11% faster. No config needed — it's the default. Other additions: selective activation checkpointing for SFT, assistant-only loss for vision datasets, better conversation logging.

Related event: TRL v1.15 Enables Fused LM Head by Default, Cutting Peak VRAM up to 82%(2 posts)→

Original post →

More from Infra

Infra channel →