Qwen2.5-72B Local Benchmark: 35 vs 65 Tokens/s Configs Analyzed
LittleCelebration412 · reddit · 2026-09-01
User compares two Qwen2.5-72B configurations on RTX 3090 (35 vs 65 tok/s). Differences stem from quantization (Q5 vs Q4), KV cache (Q8 vs Q4), concurrency slots (4 vs 1), and Multi-Token Prediction (MTP) settings, highlighting trade-offs between speed, quality, and memory.
More from Infra
- Spending $60k on Macs for Local LLMs Still Beats by $10 Cloud Subscription — leebase65 · 2026-09-01
- Call for agent infra: Who will build the open source Codex-style browser? — hwchase17 · 2026-09-01
- Why is NVIDIA MIG still limited to 7 instances on B300/Rubin? — StasBekman · 2026-09-01
- Polymarket: 69% Chance Any State Enacts Data Center Moratorium by 2026 — Polymarket · 2026-09-01
- NVIDIA Groq 3 LPX hits 3,431 tokens/s output at 100K context in benchmarks — scaling01 · 2026-09-01
- NVLink adoption deepens investment in NVIDIA hardware — BenBajarin · 2026-09-01