Same quant, same 4090: why is Qwen 3.8 slower than 3.6 locally?
Specialist-2193 · reddit · 2026-08-19
A Reddit user asks why Qwen 3.8 runs slower than 3.6 locally despite identical quantization (Q4KXL + MTP), identical hardware (RTX 4090), latest llama.cpp build, and 150k context. With the same architecture and similar MTP prediction quality, throughput should match — the user wonders what besides thinking output length could cause the slowdown. A concrete local-inference performance debugging discussion.
More from Infra
- Inference Era Competition: Cloud Decisions Driven by Revenue per Watt — BenBajarin · 2026-08-19
- Code.Storage Opens Signups: Unlimited Git Infrastructure for AI Agents — dsp_ · 2026-08-19
- Lava lamps help secure internet data — dreamwieber · 2026-08-19
- Why cloud providers stopped auctioning compute: The $999 spot instance incident — lauriewired · 2026-08-19
- Engy launches permissionless TEE worker onboarding for confidential GPU compute — markjeffrey · 2026-08-19
- Linux Desktop Share Surges Past 10%, Driven by AI Developers — DavidLinthicum · 2026-08-19