Threadripper 3975WX + 4070 Ti Super gets only 10 tok/s on Qwen 27B Q5 — tuning help wanted
ifjo · reddit · 2026-09-13
A Reddit user asks for llama.cpp tuning advice on a Threadripper 3975WX (256GB RAM) + RTX 4070 Ti Super box built for virtualization.
- Running unsloth's Qwen 27B Q5 quantization via llama.cpp server, getting 10 tokens/s
- Only flags used: --ctx-size and --jinja
- Question: is 10 tok/s the practical ceiling for this quant without a better GPU?
More from Infra
- Same model, different harness cuts cost by a third: Redditor builds live LLM cost router — leebase65 · 2026-09-13
- Why rogue agents can't pay their own compute bills: the comparative-advantage problem — voooooogel · 2026-09-13
- Databricks strips Spark tuning knobs down to one, leaving veterans wondering where to learn the internals — Zachly · 2026-09-13
- By 2030 healthcare advances will come from queryable databases — don't ban datacenters, negotiate with them — krishnan · 2026-09-13
- Apple execs' 'thermal shadow' concept: A20 packaging moves memory out of the thermal path — BenBajarin · 2026-09-13
- Community ports experimental DeepSeek V4.1 steering into antirez's C-based ds4 CLI — antirez · 2026-09-13