928-Star Wiki Details Running Qwen3.5-397B and Kimi-K2.5 on NVLink-Free PCIe RTX 6000
TheZachMueller · x · 2026-09-07
A community-maintained "RTX PRO 6000 Blackwell LLM Wiki" is trending on GitHub (928 stars), focused on serving frontier LLMs like Qwen3.5-397B, Kimi-K2.5, and GLM-5 on NVLink-free PCIe GPUs (SM120).
It goes well beyond launch snippets, offering:
- Reproducible Docker builds plus detailed vLLM and SGLang runbooks
- Benchmark tables and KLD-based quality checks
- Quantization notes and debugging of DCP/MTP/DSpark/DFlash
- PCIe topology work and a regression history
A companion Discord serves as the community workbench, making this a near-complete field manual for local deployment of very large models on PCIe-only machines.
More from Infra
- M1 64GB Mac is no match for local Astra: 'this thing is crawling' — natesiggard · 2026-09-07
- KV Cache Engineering for LLM Serving: 12 Techniques Explained With Trade-offs — AccBalanced · 2026-09-07
- SlimServe open-sources low-cost LLM serving for consumer GPUs — QuixiAI · 2026-09-07
- AI buildout debt hits $570B, now rivaling the entire US muni bond market's annual supply — ivan_bezdomny · 2026-09-07
- Prediction: compute is moving from rack-scale to datacenter-scale as bottlenecks shift outward — AccBalanced · 2026-09-07
- AI meme: turning off the tap while brushing teeth to "conserve water for the datacenter buildout" — EigenGender · 2026-09-07