Quartermaster: open-source local server runs Qwen-Image 2.1 Turbo without ComfyUI
OneMoreName1 · reddit · 2026-10-11
A developer released Quartermaster, a free MIT-licensed local inference server built on stable-diffusion.cpp that also serves LLMs, TTS and speech-to-text, swapping models in and out of VRAM. The app auto-writes configs, matches VAEs and text encoders, and sets 8 steps / cfg 1 for Qwen-Image 2.1 Turbo; it also supports real transparent PNGs via an RGBA VAE prompt. Running it requires three files: the Q80 DiT GGUF, the 2.1 VAE bf16, and Qwen3-VL-8B-Instruct GGUF plus mmproj.
More from Infra
- A 37ms Gap Left GPUs Idle ~40% of the Time; Fixing It Removed the Tiny Idles — HankYeomans · 2026-10-11
- Musk courts chip engineers as Terafab targets 10 chip designs a year — elonmusk · 2026-10-11
- RTX 5090 Hits $5,000 as AI Firms Buy Gaming GPUs by the Pallet — chemist_slime · 2026-10-11
- Qwen3.8-27B Hits 140 tok/s on a Single RTX 3090 with a CUDA Megakernel, KL Divergence 0.0009 — Adorable_Weakness_39 · 2026-10-11
- Report: Anthropic signs $11.6B, 7-year compute deal with Akamai — nikola_mr64990 · 2026-10-11
- Musk confirms Terafab will modify existing tools and build custom equipment to lift wafer output — elonmusk · 2026-10-11