Qwen3.8-27B EXL3 one-click kit brings quality local LLM to 16-32GB consumer GPUs

udmrzn · x · 2026-09-13

MiaAI Lab released a one-click serving kit for Qwen3.8-27B in turboderp's EXL3 quants, targeting consumer NVIDIA cards with 16-32GB VRAM (RTX 3090, 5060 Ti, 5070 Ti, 5090 and more). The kit auto-picks a quant that fits your VRAM (2.0bpw floor for 16GB), sets up its own Python env, downloads weights, serves an OpenAI-compatible endpoint and opens a chat UI, with identical behavior on Windows and Linux. A physician reports using it locally to build an educational ventilator simulator with an infinite random case generator in hours.

Original post →

More from Infra

Infra channel →