Squeezing Qwen 35B on an RX 6700 XT: a llama.cpp 100k-context tuning log

Loose_Doubt367 · reddit · 2026-09-03

Running unsloth Qwen3.6-35B-A3B Q4KM on an RX 6700 XT 12GB, the author keeps 100k context fixed and targets 45 tokens/sec. Two llama-server launch configs are shared: one for coding using ngram speculative decoding (match 8–32), one for general use with MTP draft decoding (p-min 0.75), plus q80 KV cache, dio load mode, and thread/batch tuning. He is asking the community for lesser-known llama.cpp options to gain more speed without hallucinations.

Original post →

More from Infra

Infra channel →