Qwen2.5-72B Local Benchmark: 35 vs 65 Tokens/s Configs Analyzed

LittleCelebration412 · reddit · 2026-09-01

User compares two Qwen2.5-72B configurations on RTX 3090 (35 vs 65 tok/s). Differences stem from quantization (Q5 vs Q4), KV cache (Q8 vs Q4), concurrency slots (4 vs 1), and Multi-Token Prediction (MTP) settings, highlighting trade-offs between speed, quality, and memory.

Original post →

More from Infra

Infra channel →