User reports 44 tok/s ingestion and 8 tok/s generation for GLM 5.2 on a $900 rig
naunen · reddit · 2026-07-27
A Reddit user reports local speed numbers for GLM 5.2 on a $900/$€ system.
- Hardware: 4x 8880 v4 CPUs, 1 TB of 32-channel DDR3 RAM, and 2x RTX 3060 12 GB.
- Model setup: 3-bit quantized version.
- Reported performance: 20k context, about 44 tokens/s ingestion and 8 tokens/s generation.
- The user asks whether these are good numbers for the cost and invites others to beat the tokens-per-dollar ratio.
More from Infra
- Users ask whether RAID 0 NVMe setups improve local large-model runs on Pulsar — Wyldkard79 · 2026-07-27
- GLM-5.2 inference on RTX 5090s jumps from 30 tok/s to 80–110 tok/s — markjeffrey · 2026-07-27
- Hermes Agent says progressive tool disclosure scales MCP tools with near-zero accuracy loss — Teknium · 2026-07-27
- Laguna tests 2.75 and 3.25 bpw quantization with NVFP4 experts and FP8 KV cache — QuixiAI · 2026-07-27
- Chutes says it trained a 20B model for under $10 an hour using rented GPUs across two continents — markjeffrey · 2026-07-27
- llama.cpp merges support for Minimax M3 with MSA — Time_Reaper · 2026-07-27