Optimal Settings for Llama.cpp + Qwen 3.8: n-max 4 Fastest

GodComplecs · reddit · 2026-08-31

Testing on RTX 3090 with the Qwen 3.8 UD Q4 KM model revealed that n-max 4 combined with spec draft min p 0.7 yields the fastest speed for harness workflows. The configuration supports 205k context with Flash Attention enabled. Benchmarks showed a generation speed of 70 tks in testing, outperforming previous mtp2 settings.

Original post →

More from coding & agent

coding & agent channel →