Qwen3.8 Benchmarks: MTP Settings Impact Throughput, Q4 Outperforms Q8

New-Inspection7034 · reddit · 2026-08-17

Author tested Qwen3.8-27B (GGUF) on an RTX Pro 6000 Blackwell with an agent harness, revealing counterintuitive findings on performance tuning and quantization:

1. MTP Parameters Do Not Transfer

2. Q4 Quantization Beats Q8

3. reasoningeffort > Quantization

4. Model Upgrade Exposed Tooling Bugs

Original post →

More from Infra

Infra channel →