Qwen3.8-27B EXL3 one-click kit brings quality local LLM to 16-32GB consumer GPUs
udmrzn · x · 2026-09-13
MiaAI Lab released a one-click serving kit for Qwen3.8-27B in turboderp's EXL3 quants, targeting consumer NVIDIA cards with 16-32GB VRAM (RTX 3090, 5060 Ti, 5070 Ti, 5090 and more). The kit auto-picks a quant that fits your VRAM (2.0bpw floor for 16GB), sets up its own Python env, downloads weights, serves an OpenAI-compatible endpoint and opens a chat UI, with identical behavior on Windows and Linux. A physician reports using it locally to build an educational ventilator simulator with an infinite random case generator in hours.
More from Infra
- Community ports experimental DeepSeek V4.1 steering into antirez's C-based ds4 CLI — antirez · 2026-09-13
- Threadripper 3975WX + 4070 Ti Super gets only 10 tok/s on Qwen 27B Q5 — tuning help wanted — ifjo · 2026-09-13
- MHA, MQA, GQA and MLA explained by what happens to the K/V cache during decoding — techNmak · 2026-09-13
- Early OpenAI employee says 'winning' AGI is outdated — 99% of future compute will run locally — GregCook2011 · 2026-09-13
- Benchmarks show ComfyUI in Docker (CUDA 12.4) runs at 0% penalty if you fix the /dev/shm OOM crash — fluxdraw · 2026-09-13
- Nex-N2.5-mini-MLX-4bit hits 133.6 tok/s on Apple M5 Max — DerTomsn · 2026-09-13