Deploying GLM-5.3 Flash on 2x RTX PRO 6000: Quantization and vLLM compatibility
No-Paper-557 · reddit · 2026-08-28
A user discusses the best setup for running GLM-5.3 Flash on dual RTX PRO 6000 96GB cards (192GB total VRAM). The debate centers on choosing the current IQ4XS GGUF quantization versus waiting for a better-fitting NVFP4 quant. The post also inquires about vLLM compatibility issues with NVFP4 on SM120, specifically regarding the sparse attention/NoPE path blocker on RTX PRO 6000 Blackwell and potential workarounds.
More from Infra
- Qwen 350K Context Tested on M5 Max: Performance and Quality — Artistic_Okra7288 · 2026-08-30
- Azure Linux 4.0 Desktop Concept: PowerShell, Edge, and Copilot Pre-installed — unixterminal · 2026-08-30
- Jensen Huang: Built GPU tech first, found endless problems from graphics to molecular dynamics — r0ck3t23 · 2026-08-30
- How to build an LLM inference engine from scratch: 5-layer architecture — glenbeer · 2026-08-30
- Huaqin expects super node revenue to exceed 10B RMB in 2H 2026 — zephyr_z9 · 2026-08-30
- Nvidia is generating $1 billion a day, a business scale deemed absurd years ago — shauntrennery · 2026-08-30