Benchmarking low-thinking Qwen quants: 26-33% fewer tokens, ThinkingCap runs 23% faster on a 7900 XTX

DerTomsn · reddit · 2026-09-25

The author benchmarked two "less thinking" Qwen quants (Swift and ThinkingCap) against a regular unsloth Q4KM quant on a single 7900 XTX, across 4 scenarios with 2 runs each.

Key findings:

Caveats: only 2 runs per quant, so variance matters; the low ThinkingCap prefill speed needs more investigation. A side-by-side comparison is available on llm-bench.io.

Original post →

More from Infra

Infra channel →