Project Maya runs 321B GLM-5.3-Flash locally on two old V100s at up to 40 tok/s

lxfater · x · 2026-10-07

Open-source Project Maya runs GLM-5.3-Flash — a 321B-parameter MoE (about 18B active per token, up to 1M context) — on modest local hardware. Author lxfater reports 2× V100 32GB + 32GB RAM + NVMe SSD hitting up to 40 token/s. The trick: hot experts live in VRAM while the rest are paged in from RAM/SSD on demand, with a 90GB Maya-S quantized model by default. It ships a web chat UI, image understanding, and OpenAI/Anthropic-compatible APIs for coding agents, built on the Strata engine (which previously ran Qwen's 100B-class models on 12GB VRAM + 64GB RAM). MIT-licensed, Linux-only for now (no Windows/WSL2).

Original post →

More from Infra

Infra channel →