QuixiAI open-sources SlimServe, the inference stack behind the 8x3090 Qwen deployment

QuixiAI · x · 2026-09-06

Eric Hartford open-sourced SlimServe on GitHub, the inference stack behind his 8x RTX 3090 deployment of Qwen3.8-Flash-Next. The repo (19,588 commits) includes vLLM integration, benchmark scripts, a perf worklog, the 3090 deployment blog writeup, plus engineering docs on GLM5.2 pipeline tracing, Metal scaling and quantization notes. Developers can reproduce the setup via the blog doc and benchmarks directory.

Original post →

More from Infra

Infra channel →