SlimServe open-sources low-cost LLM serving for consumer GPUs
QuixiAI · x · 2026-09-07
Developer QuixiAI open-sourced SlimServe, a project for serving quantized LLMs on consumer GPUs like the RTX 3090. The repo includes deployment guides for Qwen3 8-bit and GLM-5.2, benchmark scripts, quantization notes, Metal scaling tests, and AMD Marlin GPTQ experiments — a useful engineering reference for low-cost local inference.
More from Infra
- Jensen Huang confirms GPT-6 Astra trained on 100K+ Grace Blackwell NVL72 — himanshustwts · 2026-09-07
- Wan2GP lands on Pinokio: one-click AI video generation for 6GB+ VRAM machines — cocktailpeanut · 2026-09-07
- Wan2GP AMD edition hits Pinokio, supporting all RDNA 2-4 discrete GPUs — cocktailpeanut · 2026-09-07
- Buying a room full of hardware to run OpenClaw as supreme rage bait — HankYeomans · 2026-09-07
- Google says high-performance memory now exceeds 75% of an AI server's bill of materials — Beth_Kindig · 2026-09-07
- Fully automated product demo videos: local LLMs, 34 episodes, zero human editing — Ok_Cartographer_6086 · 2026-09-07