PreFT Paper Accepted at NeurIPS: Prefill-Only LoRA Adapters Speed Up Multi-Adapter Serving

aryaman2020 · x · 2026-09-25

A new NeurIPS paper introduces Prefill-Only Fine Tuning (PreFT): since prefill and decode are very different inference workloads, serving many LoRA adapters at once slows decode badly due to memory-bound bottlenecks. PreFT trains and applies adapters only at prefill, eliminating them at decode and speeding up multi-adapter serving with limited performance loss.

Original post →

More from Infra

Infra channel →