vLLM Introduces Experimental AFD Plugin to Boost MoE Inference
vLLM released an experimental AFD plugin that decouples attention and FFN layers in MoE inference. Tests show this approach improves throughput by 11.3% and reduces TTFT by 47%.
2026-07-24 ~ 2026-07-24 · 2 related posts
- vLLM adds an experimental plugin to split attention and FFN serving — vllm_project · 2026-07-24
- vLLM AFD tests show 11.3% higher decode throughput and 47% lower TTFT — vllm_project · 2026-07-24