Squeezing Qwen 35B on an RX 6700 XT: a llama.cpp 100k-context tuning log
Loose_Doubt367 · reddit · 2026-09-03
Running unsloth Qwen3.6-35B-A3B Q4KM on an RX 6700 XT 12GB, the author keeps 100k context fixed and targets 45 tokens/sec. Two llama-server launch configs are shared: one for coding using ngram speculative decoding (match 8–32), one for general use with MTP draft decoding (p-min 0.75), plus q80 KV cache, dio load mode, and thread/batch tuning. He is asking the community for lesser-known llama.cpp options to gain more speed without hallucinations.
More from Infra
- Nvidia's vertical integration signals next-gen app layer push beyond CUDA — RachelVT42 · 2026-09-03
- ARK mid-year AI review: compute buildout and adoption keep beating expectations — downingARK · 2026-09-03
- Buying a refurbished 8xA100 server to colocate and rent out: one user's GPU math — Alarming-Ad8154 · 2026-09-03
- Hidden consumable under AI servers: PCB drill-bit demand is growing far faster than board volumes — tengyanAI · 2026-09-03
- Vera Rubin NVL72 hits up to 10X tokens per MW vs GB200 on DeepSeek R1, CoreWeave data shows — Beth_Kindig · 2026-09-03
- Tencent Hunyuan 770B compressed from ~1.5TB to ~214GiB with mixed quantization — Aiden_Tech_Ai · 2026-09-03