Same quant, same 4090: why is Qwen 3.8 slower than 3.6 locally?

Specialist-2193 · reddit · 2026-08-19

A Reddit user asks why Qwen 3.8 runs slower than 3.6 locally despite identical quantization (Q4KXL + MTP), identical hardware (RTX 4090), latest llama.cpp build, and 150k context. With the same architecture and similar MTP prediction quality, throughput should match — the user wonders what besides thinking output length could cause the slowdown. A concrete local-inference performance debugging discussion.

Original post →

More from Infra

Infra channel →