Qwen3.8-27B Serving Configs: DGX Spark vLLM and RTX 4090 llama.cpp

erdaltoprak · reddit · 2026-08-15

Reddit user erdaltoprak shares serving configurations for Qwen3.8-27B: one for Nvidia DGX Spark with vLLM (NVFP4, 262k context, MTP speculative decoding), and one for RTX 4090 with llama.cpp (GGUF Q4KM, 131k context, flash attention). Docker Compose files provided.

Original post →

More from coding & agent

coding & agent channel →