vLLM Introduces Experimental AFD Plugin to Boost MoE Inference

vLLM released an experimental AFD plugin that decouples attention and FFN layers in MoE inference. Tests show this approach improves throughput by 11.3% and reduces TTFT by 47%.

2026-07-24 ~ 2026-07-24 · 2 related posts