Full guide: running vLLM on Windows with Docker, WSL2, and an RTX PRO 6000

Demonicated · reddit · 2026-08-19

Windows users wanting vLLM usually settle for simpler tools like LM Studio; the author worked out a complete flow on an RTX PRO 6000 Blackwell (96GB) and let AI help write the guide. It runs Qwen3.8-27B as an OpenAI-compatible vLLM server via Windows 11 + WSL 2 + Docker Desktop, supporting vision, reasoning, tool calling, prefix caching, and MTP speculative decoding — with concurrent agent interactions at little throughput cost.

Key gotchas:

Original post →

More from Infra

Infra channel →