A modded RTX 3090 with 48GB RAM runs a 27B Qwen GGUF locally via SlimServe

MaziyarPanahi · x · 2026-08-20

QuixiAI showed off a local inference setup: an RTX 3090 modded to 48GB of RAM running UnslothAI's Dynamic Qwen3.8-27B Q3KXL GGUF quantized model with TurboQuant and DFlash 2, using their own SlimServe engine. Maziyar Panahi asked about the actual tokens/s throughput — a sign that single-GPU, high-VRAM local deployment of mid-size models remains a hot topic.

Related event: Modded 48GB RTX 3090 Runs 27B Model Locally at High Speeds(4 posts)→

Original post →

More from Infra

Infra channel →