GLM 5.2 Hits 385 tok/s on 8x RTX PRO 6000 with NVFP4
max_paperclips · x · 2026-08-08
A developer successfully ran the GLM 5.2 model on 8x RTX PRO 6000 GPUs using the NVFP4 quantization format, achieving an inference speed of 385 tokens/s.
The poster noted that initial runs were buggy and incredibly slow, taking over a week of debugging with the help of AI agents to reach this performance milestone.
More from Infra
- Nvidia to Invest Up to $3 Billion in Blackstone-Backed Power Firm — pstAsiatech · 2026-08-08
- AURORA-LM: A 1B Continuous Diffusion Language Model Trained on Ascend NPU — 机器之心 · 2026-08-08
- Local Video Generation: MiniMax Lags Far Behind LTX in Inference Speed — PhilosopherSweaty826 · 2026-08-08
- Deep Dive: How Weak is the Evidence for China's Role in US Data Center Backlash? — AndyMasley · 2026-08-08
- $50k Bet Challenges SemiAnalysis on SpaceX AI Compute and ARR Forecasts — generativist · 2026-08-08
- Data Center Boom Drives 'Second Great Construction Divergence' and Job Growth — ivan_bezdomny · 2026-08-08