Qwen2.5-72B INT8 benchmark: vLLM config and optimization on A40

OvertaxedOne · reddit · 2026-08-26

A Reddit user shared their experience running Qwen2.5-72B INT8 on an A40 GPU, achieving 25 TPS but hitting memory limits at 256K context (KV Cache FP8). The user is considering switching to INT4 for better speed/headroom and asks about the quality drop, specifically for tool calling with Hermes.

vLLM Config Highlights:

Original post →

More from Infra

Infra channel →