Modded 48GB RTX 3090 Runs 27B Model Locally at High Speeds
Developer QuixiAI showed a local setup pairing a modded 48GB RTX 3090 with the SlimServe engine to run quantized Qwen models, and plans optimized Unsloth AI kernels targeting up to 300 tok/s, drawing wide community attention.
2026-08-20 ~ 2026-08-20 · 4 related posts
- Baby's first PC: 3090 with local LLM inference — QuixiAI · 2026-08-20
- Building a local inference rig: RTX 3090 with 48GB RAM running Qwen3.8-27B — MaziyarPanahi · 2026-08-20
- A modded RTX 3090 with 48GB RAM runs a 27B Qwen GGUF locally via SlimServe — MaziyarPanahi · 2026-08-20
- Unsloth AI to launch new kernels, predicting 300 tok/s on C8 — QuixiAI · 2026-08-20