GLM-5.2 hits 12.2 tok/s on 16 AMD MI50s with llama.cpp RPC
Legal-Ad-3901 · reddit · 2026-07-23
- A user reports GLM-5.2 UD-Q4KXL GGUF running on a 16× AMD MI50 32GB setup via llama.cpp RPC.
- On a real 10.7k-token document, it reached 12.2 tok/s output and 30.9 tok/s input; with two parallel requests, aggregate throughput was 14.5 tok/s.
- At 10.7k context, output stayed at 10.2 tok/s with coherent long-form generation.
- The cluster uses 2× 8-GPU nodes, 512 GB total VRAM, 100W cap per card, and 10 GbE DAC interconnect.
- The author says earlier attempts with a custom vLLM-gfx906 build hit decent speed but degraded past 10k context, and they hope someone can get the Moby Dick repo working better with it.
More from Infra
- Report says Mistral’s sovereign AI stack still runs on Microsoft infrastructure — shashib · 2026-07-23
- Analyst: AI Data Centers Are Beginning to Resemble Semiconductor Fabs — BenBajarin · 2026-07-23
- AMD’s MI300X and MI350 win praise as Anthropic’s reported MI450 deal signals a CUDA challenge — tzmartin · 2026-07-23
- Token speed fails to predict real task time or cost, a 13-model probe finds — According-Floor5177 · 2026-07-23
- FiveClaw adds a managed MCP codespace for FiveM AI development — nytro_Haze · 2026-07-23
- Open-source QuantProbe predicts local-LLM speed on a 2016 PC and a GTX 1060 — Ok_Brush_3449 · 2026-07-23