Deploying GLM-5.3 Flash on 2x RTX PRO 6000: Quantization and vLLM compatibility

No-Paper-557 · reddit · 2026-08-28

A user discusses the best setup for running GLM-5.3 Flash on dual RTX PRO 6000 96GB cards (192GB total VRAM). The debate centers on choosing the current IQ4XS GGUF quantization versus waiting for a better-fitting NVFP4 quant. The post also inquires about vLLM compatibility issues with NVFP4 on SM120, specifically regarding the sparse attention/NoPE path blocker on RTX PRO 6000 Blackwell and potential workarounds.

Original post →

More from Infra

Infra channel →