QuixiAI open-sources SlimServe, the inference stack behind the 8x3090 Qwen deployment
QuixiAI · x · 2026-09-06
Eric Hartford open-sourced SlimServe on GitHub, the inference stack behind his 8x RTX 3090 deployment of Qwen3.8-Flash-Next. The repo (19,588 commits) includes vLLM integration, benchmark scripts, a perf worklog, the 3090 deployment blog writeup, plus engineering docs on GLM5.2 pipeline tracing, Metal scaling and quantization notes. Developers can reproduce the setup via the blog doc and benchmarks directory.
More from Infra
- NanoDiffuser: sub-400 MiB text-to-image model running in-browser via WebGPU — jason_mayes · 2026-09-06
- Wall Street veteran: the world's biggest data center consumes water like just 3 golf courses — rohanpaul_ai · 2026-09-06
- Japan says $550B U.S. investment pact advancing, AI and semiconductors to play major role — Polymarket · 2026-09-06
- Serving Qwen3.8-Flash-Next at 262K context on 8 RTX 3090s hits 1,200 tok/s — QuixiAI · 2026-09-06
- T-Glass shortage worsens: Kinsus losing 10-15% of monthly ABF revenue, 25% capacity expansion planned for 2027 — zephyr_z9 · 2026-09-06
- Bump-less 3D stacking goes practical: Intel Diamond Rapids first, AMD Zen rumored next — bookwormengr · 2026-09-06