Qwen2.5-72B INT8 benchmark: vLLM config and optimization on A40
OvertaxedOne · reddit · 2026-08-26
A Reddit user shared their experience running Qwen2.5-72B INT8 on an A40 GPU, achieving 25 TPS but hitting memory limits at 256K context (KV Cache FP8). The user is considering switching to INT4 for better speed/headroom and asks about the quality drop, specifically for tool calling with Hermes.
vLLM Config Highlights:
- Model: lued/Qwen2.5-72B-INT8-W8A16-MTP
- Key flags: --attention-backend FLASHINFER, --quantization compressed-tensors, --kv-cache-dtype fp8e4m3, --speculative-config {"method":"mtp","numspeculativetokens":1}
- Enabled prefix caching, chunked prefill, and auto tool choice.
More from Infra
- Insider calls Apple's new lineup "serious beef" — adrianscottcom · 2026-08-26
- Latency reduction in MoE models comes at high cost — firstadopter · 2026-08-26
- Applied Compute launches AC2 to enable teams to train and serve custom frontier models — ypatil125 · 2026-08-26
- CBRS 3D DRAM Wafer Estimated at 1TB Raw Capacity — BenBajarin · 2026-08-26
- BART Spends $65M/Year on Electricity, 1/700th of CA Usage — cis_female · 2026-08-26
- CXMT's HBM Yield at 25%; SK Hynix and Micron Maintain High-End Lead — demian_ai · 2026-08-26