Qwen3.6-35B Runs Smoothly on a Single RTX 4070
Developers report Qwen3.6-35B (MoE, 3B active parameters) runs well on a single RTX 4070 after quantization, with full llama-server configuration details shared for local deployment.
2026-09-22 ~ 2026-09-22 · 3 related posts
- Qwen3.6-35B runs surprisingly well on a single RTX 4070, dev reports — haydendevs · 2026-09-22
- Qwen3.6-35B-A3B runs well on an RTX 4070, dev reports with UD-Q4_K_XL quant — haydendevs · 2026-09-22
- Full llama-server config for local Qwen3.6-35B-A3B with MTP speculative decoding — haydendevs · 2026-09-22