Running GLM-5.2 and Kimi K3 Locally on RTX 5090: Hits 129 tok/s
markjeffrey · x · 2026-07-31
A user shared local inference benchmarks using a half-price RTX 5090 consumer setup.
- GLM-5.2: Achieved a decode speed of 129 tok/s.
- Kimi K3: Reached 108 tok/s.
- Context: This test responds to a previous frontier deployment by the tinygrad team achieving 120 tok/s for GLM-5.2 on AMD boxes, proving that consumer-grade GPUs offer highly cost-effective performance for running large models.
More from Infra
- Running Inkling Small on a Single Node: Native Voice Interaction Under 500ms — andimarafioti · 2026-07-31
- LLM Inference Costs Plunge: Token Prices Drop to 1/13th in Four Months — charliermarsh · 2026-07-31
- Samsung Earnings: Agentic AI Drives Surge in Enterprise SSD Demand — davidyin44 · 2026-07-31
- vLLM Releases Inkling-Small Deployment Guide: Runs on Minimum 180GB VRAM — vllm_project · 2026-07-31
- Malaysian Activists Successfully Halt Data Centre Construction — im_mansigupta · 2026-07-31
- Meta's AI Infrastructure Lease Obligations Surge 53% in Three Months to Nearly $279 Billion — Polymarket · 2026-07-31