Running Qwen3.8-27B-NVFP4 locally with vLLM at 1M context: full config shared
No_Night679 · reddit · 2026-09-16
A self-described beginner shared a working native (non-container) vLLM setup serving unsloth/Qwen3.8-27B-NVFP4 with a 1M-token context window, posting the full systemd unit.
Key settings: 4-way tensor parallelism with gpu-memory-utilization 0.91, fp8 KV cache, MTP speculative decoding with 3 speculative tokens, and hf-overrides injecting YaRN RoPE parameters (ropetheta 1e7, factor 4.0, originalmaxpositionembeddings 262144) to stretch context to 1,000,000 tokens. Qwen3 reasoning and qwen3xml tool-call parsers plus auto tool choice are enabled.
The author reports it runs well so far but is unsure whether the config can be tuned further — a directly copyable starting point for anyone attempting ultra-long context on local hardware.
More from Infra
- Periodic Labs: specialized models beat GPT-6 on science evals with 1,300 H200s — vwxyzjn · 2026-09-16
- E2B Renames Its Open-Source Project to E2B Runtime, the Backend Behind Its AI Agent Cloud — badphilosopher · 2026-09-16
- NVIDIA and Palantir plant a flag for sovereign AI, starting with NVIDIA — postalex · 2026-09-16
- Ben Bajarin bullish on Credo: DustPhotonics deal and vertical integration undervalued by market — BenBajarin · 2026-09-16
- Unconventional AI shows analog oscillator chip, claims 1000x power efficiency over Nvidia in 2 years — PTrubey · 2026-09-16
- RAM-Pooling Across 3 Old Devices Runs a 40B Model at 16 tok/s — With a Counterintuitive CPU Finding — Medicine_Blogscanner · 2026-09-16