Qwen3.8-27B Serving Configs: DGX Spark vLLM and RTX 4090 llama.cpp
erdaltoprak · reddit · 2026-08-15
Reddit user erdaltoprak shares serving configurations for Qwen3.8-27B: one for Nvidia DGX Spark with vLLM (NVFP4, 262k context, MTP speculative decoding), and one for RTX 4090 with llama.cpp (GGUF Q4KM, 131k context, flash attention). Docker Compose files provided.
More from coding & agent
- Perplexity launches Search SDK for agents, enabling deep research in code — AravSrinivas · 2026-08-15
- Self-bench: Open-source tool to auto-build private coding-agent benchmarks from your repo — ricklamers · 2026-08-15
- 27B agent Faraday beats Claude Opus 4.8 and GPT-5.5 on research replication via new Replica method — omarsar0 · 2026-08-15
- Seeking LiteLLM Alternatives with Enterprise-Grade Security and RBAC — Screwtapeworm69 · 2026-08-15
- What Hidden States Should an AI Agent Track When Diagnosing CI Failures? A Developer Seeks Feedback — Elegant_Quantity_583 · 2026-08-15
- AI analysis tool runs 67 minutes, gathers 1,063 evidence items, verifies 532 claims — ChrisGPT · 2026-08-15