vLLM Releases Experimental AFD Plugin to Boost MoE Inference
vLLM introduced an experimental AFD plugin that decouples attention and FFN serving for MoE inference. Tests show the plugin improves throughput by 11.3% and reduces time-to-first-token (TTFT) by 47%.
2026-07-24 ~ 2026-07-25 · 3 related posts
- vLLM adds an experimental plugin to split attention and FFN serving — vllm_project · 2026-07-24
- vLLM AFD tests show 11.3% higher decode throughput and 47% lower TTFT — vllm_project · 2026-07-24
1 near-duplicate retellings: AccBalanced