GLM 5.3 Flash runs locally on dual V100s: 93GB MoE weights at ~20 tokens/s
lxfater · x · 2026-10-06
A user appears to have gotten a suspected GLM 5.3 Flash checkpoint running locally before any official release, and a reposter notes it runs for them too.
- Model file is 93.1GB, the smallest available so far, likely MoE (author shared expert placement as proof)
- No MTP support yet; decode speed around 20 tokens/s
- Surprisingly modest hardware: 2x Tesla V100 32GB, Xeon E5-2690 V4, 32GB DDR4 quad-channel, NVMe 3.0
- Runs on a custom version of the Strata inference engine; the author plans further optimization and eventual release
More from Models
- Opus 5.5 Builds a Fully Animated Pixel-Art Scene in TypeScript in ~10 Minutes — iamfakhrealam · 2026-10-06
- DINOv3 + SAM Local Vision Pipeline Counts Surgical Instruments in Real Time — MaziyarPanahi · 2026-10-06
- Cohere details Tiny Aya L2-Thinker: data mixing unlocks reasoning in 60 languages — Cohere_Labs · 2026-10-06
- 21 hacks to avoid hitting Claude's usage limits — aitrendz_xyz · 2026-10-06
- Low Effort Can Burn More Tokens: Qwen3.8-27B-pi Fine-Tunes Coding Agent Effort Ordering — lmoroney · 2026-10-06
- Qwen 4 reportedly planned for end of October, per alleged Alibaba 0-day partner — Dependent_Hunter_155 · 2026-10-06