A modded RTX 3090 with 48GB RAM runs a 27B Qwen GGUF locally via SlimServe
MaziyarPanahi · x · 2026-08-20
QuixiAI showed off a local inference setup: an RTX 3090 modded to 48GB of RAM running UnslothAI's Dynamic Qwen3.8-27B Q3KXL GGUF quantized model with TurboQuant and DFlash 2, using their own SlimServe engine. Maziyar Panahi asked about the actual tokens/s throughput — a sign that single-GPU, high-VRAM local deployment of mid-size models remains a hot topic.
Related event: Modded 48GB RTX 3090 Runs 27B Model Locally at High Speeds(4 posts)→
More from Infra
- Report: Stripe to buy OpenRouter for ~$7.5B, controlling AI unit economics — rohanpaul_ai · 2026-08-20
- DeepSpace SDK aims to bridge prototype-to-product gap — JaynitMakwana · 2026-08-20
- Supabase increases Edge Functions limits for Pro and Teams — dshukertjr · 2026-08-20
- Codon Compiler Boosts Python Performance 10-100x While Retaining Library Access — KhuyenTran16 · 2026-08-20
- UK AI Chip Startup CallosumAI Raises $100M Seed — HZoete · 2026-08-20
- UK Chip Startups Raise Over $900M in Two Weeks — HZoete · 2026-08-20