Benchmarking low-thinking Qwen quants: 26-33% fewer tokens, ThinkingCap runs 23% faster on a 7900 XTX
DerTomsn · reddit · 2026-09-25
The author benchmarked two "less thinking" Qwen quants (Swift and ThinkingCap) against a regular unsloth Q4KM quant on a single 7900 XTX, across 4 scenarios with 2 runs each.
Key findings:
- Token savings are real: total tokens drop from 66k (base) to 49k (ThinkingCap, -26%) and 45k (Swift, -33%).
- Different trade-offs: ThinkingCap's prefill is oddly slow (100 t/s, TTFT 6-11s) with barely-changed decode (43 t/s); Swift has the fastest prefill (614 t/s, 1s TTFT) but decode drops 33% to 32 t/s.
- Quality holds: eval scores stay in the 84-87 band, same as base.
- Runtime: ThinkingCap finishes 23% faster than base (1087s vs 1407s) — token savings outweigh the slow prefill; Swift is break-even (1400s) due to slower decode.
Caveats: only 2 runs per quant, so variance matters; the low ThinkingCap prefill speed needs more investigation. A side-by-side comparison is available on llm-bench.io.
More from Infra
- 320Gb/s unamplified transmission demoed with 100GHz Ge PD and TFLN MZM on foundry SiPh — jwt0625 · 2026-09-26
- Comparing Ghent and imec silicon photonics PDs: 320 Gb/s links but 5 dB grating coupler loss — jwt0625 · 2026-09-26
- House votes 417-3 on bill making data centers pay added grid costs — VraserX · 2026-09-26
- Railway launches free no-account VMs (59 min) and OpenCode cloud agents — jasonkneen · 2026-09-26
- DistribAI v2 pools free Colab/Kaggle GPUs into ~2x RTX 5090 training power — Enderchef · 2026-09-26
- steipete pushes back on agent cost concerns: capable models now $0.1 per million tokens — steipete · 2026-09-26