Full llama-server config for local Qwen3.6-35B-A3B with MTP speculative decoding

haydendevs · x · 2026-09-22

The author shares a complete llama-server launch command for running Qwen3.6-35B-A3B-MTP-UD-Q4KXL locally:

Serving on 0.0.0.0:8080 as an OpenAI-compatible API — a ready-to-copy recipe for local MoE deployment.

Related event: Qwen3.6-35B Runs Smoothly on a Single RTX 4070(3 posts)→

Original post →

More from Infra

Infra channel →