Use Flashinfer for VLLM on Ampere Hardware

mayo551 · reddit · 2026-08-17

Through trial and error, a user found that Flashinfer is the best attention backend for VLLM on Ampere hardware. Unlike FA2 and tritonattn, Flashinfer maintains speed at high context lengths (>70k) and works seamlessly with MTP. It achieves 100+ T/S on a 4x3090 setup with Qwen 3.8 27B. Users should verify the backend in the logs and use --language-model-only for Gemma models.

Original post →

More from Infra

Infra channel →