Building a local inference rig: RTX 3090 with 48GB RAM running Qwen3.8-27B
MaziyarPanahi · x · 2026-08-20
A developer shared a setup featuring an RTX 3090 with 48GB of RAM, running a locally hosted Qwen3.8-27B model (Q3KXL quantization) using the SlimServe inference engine with UnslothAI Dynamic 3.0, TurboQuant, and DFlash 2. It highlights a specific hardware configuration for high-performance local LLM inference.
Related event: Modded 48GB RTX 3090 Runs 27B Model Locally at High Speeds(4 posts)→
More from Infra
- Qwen 3.5 9B + DFlash hits 75 tok/s on R9700: Full Setup Guide — karmakaze1 · 2026-08-20
- DeepSpace SDK aims to bridge prototype-to-product gap — JaynitMakwana · 2026-08-20
- Supabase increases Edge Functions limits for Pro and Teams — dshukertjr · 2026-08-20
- Codon Compiler Boosts Python Performance 10-100x While Retaining Library Access — KhuyenTran16 · 2026-08-20
- UK AI Chip Startup CallosumAI Raises $100M Seed — HZoete · 2026-08-20
- UK Chip Startups Raise Over $900M in Two Weeks — HZoete · 2026-08-20