SlimServe open-sources low-cost LLM serving for consumer GPUs

QuixiAI · x · 2026-09-07

Developer QuixiAI open-sourced SlimServe, a project for serving quantized LLMs on consumer GPUs like the RTX 3090. The repo includes deployment guides for Qwen3 8-bit and GLM-5.2, benchmark scripts, quantization notes, Metal scaling tests, and AMD Marlin GPTQ experiments — a useful engineering reference for low-cost local inference.

Related event: Developer forks NVIDIA open driver to unlock P2P on GeForce cards, runs 262K-context inference on 8 RTX 3090s(8 posts)→

Original post →

More from Infra

Infra channel →