GLM 5.3 Flash spotted hitting 227 tok/s with 4 concurrent streams on DGX Spark
natesiggard · x · 2026-10-08
MiaAIlab demoed GLM 5.3 Flash running locally on an NVIDIA DGX Spark, reaching roughly 227 tokens/s across 4 concurrent mixed streams — a speed the poster calls "unreal." The version number suggests an unreleased GLM 5.3 Flash variant; details remain unconfirmed.
More from Infra
- Cloud company books Delta Forge Two for 15 years at ~$5B before a single customer — YvesMulkers · 2026-10-09
- PyTorch Conference to Feature Meta's TorchTPU Cross-Hardware Portability Keynote — PyTorch · 2026-10-09
- Open-source inference cloud operator: compute demand will be desperately short if workloads shift to open models — gharik · 2026-10-08
- Andy Pavlo: AI agents create 80% of new databases and keep deleting production ones — mattturck · 2026-10-08
- San Francisco unanimously approves temporary ban on new data centers — Polymarket · 2026-10-08
- Genesis Mission partners pledge $2.4B in compute; NSF and DOE add $100M — AllThingsApx · 2026-10-08