Building a local inference rig: RTX 3090 with 48GB RAM running Qwen3.8-27B

MaziyarPanahi · x · 2026-08-20

A developer shared a setup featuring an RTX 3090 with 48GB of RAM, running a locally hosted Qwen3.8-27B model (Q3KXL quantization) using the SlimServe inference engine with UnslothAI Dynamic 3.0, TurboQuant, and DFlash 2. It highlights a specific hardware configuration for high-performance local LLM inference.

Related event: Modded 48GB RTX 3090 Runs 27B Model Locally at High Speeds(4 posts)→

Original post →

More from Infra

Infra channel →