Use Flashinfer for VLLM on Ampere Hardware
mayo551 · reddit · 2026-08-17
Through trial and error, a user found that Flashinfer is the best attention backend for VLLM on Ampere hardware. Unlike FA2 and tritonattn, Flashinfer maintains speed at high context lengths (>70k) and works seamlessly with MTP. It achieves 100+ T/S on a 4x3090 setup with Qwen 3.8 27B. Users should verify the backend in the logs and use --language-model-only for Gemma models.
More from Infra
- OpenAI Signs Record Ohio Data Center Lease with $105B Nvidia Backing — The Decoder · 2026-08-17
- Cerebras scales supply chain: 600MW data center capacity under contract, 10x manufacturing boost planned — Sethwinterroth · 2026-08-17
- SK Hynix to boost Dalian fab output by 50% by 2027 — Beth_Kindig · 2026-08-17
- Cursor's AWS VM cost bottleneck: $300-$900/month for 10 bots — Daniel_Farinax · 2026-08-17
- QVM adds ultra low-latency desktop streaming with cross-platform support — OwariDa · 2026-08-17
- Groq Raises $350M at $3.5B Valuation After Nvidia Deal — dinabass · 2026-08-17