PreFT Paper Accepted at NeurIPS: Prefill-Only LoRA Adapters Speed Up Multi-Adapter Serving
aryaman2020 · x · 2026-09-25
A new NeurIPS paper introduces Prefill-Only Fine Tuning (PreFT): since prefill and decode are very different inference workloads, serving many LoRA adapters at once slows decode badly due to memory-bound bottlenecks. PreFT trains and applies adapters only at prefill, eliminating them at decode and speeding up multi-adapter serving with limited performance loss.
More from Infra
- AI data center talent war: electrician pay up 3-5x in Texas and Louisiana, says SemiAnalysis — johncoogan · 2026-09-25
- CoreWeave turns away $100M customers until May as Nebius climbs to Platinum GPU cloud tier — johncoogan · 2026-09-25
- Inside Quail: custom vLLM scheduler, workload-aware KV cache for 1B tok/min — sh_reya · 2026-09-25
- Anthropic's CI job volume grew 25x in six months — here's how they scaled test selection — JeremyCMorgan · 2026-09-25
- Chip startups like Cerebras and Groq are becoming the next wave of neoclouds, says SemiAnalysis — johncoogan · 2026-09-25
- Anthropic Commits $11.6B Over 7 Years to Akamai Cloud, Validating Distributed Inference Thesis — pdamodaran · 2026-09-25