Project Maya runs 321B GLM-5.3-Flash locally on two old V100s at up to 40 tok/s
lxfater · x · 2026-10-07
Open-source Project Maya runs GLM-5.3-Flash — a 321B-parameter MoE (about 18B active per token, up to 1M context) — on modest local hardware. Author lxfater reports 2× V100 32GB + 32GB RAM + NVMe SSD hitting up to 40 token/s. The trick: hot experts live in VRAM while the rest are paged in from RAM/SSD on demand, with a 90GB Maya-S quantized model by default. It ships a web chat UI, image understanding, and OpenAI/Anthropic-compatible APIs for coding agents, built on the Strata engine (which previously ran Qwen's 100B-class models on 12GB VRAM + 64GB RAM). MIT-licensed, Linux-only for now (no Windows/WSL2).
More from Infra
- 土豆地里长出 token:乌兰察布已落地超 100 个 AI 数据中心项目 — pstAsiatech · 2026-10-08
- Hybrid agent pattern: cloud Gemini plans, local Gemma swarm runs 97% of tokens offline — clmt · 2026-10-08
- Nvidia-Backed Lambda Raising Up to $4B at $14.5B Valuation Ahead of 2027 IPO — darian314 · 2026-10-08
- Strata 0.1.40.1 leaves 10.7GB VRAM unused on purpose: 24% fewer cached experts, faster decode — Critical-Entry3377 · 2026-10-07
- Flama 2.0: 8-year-old Python framework now packages LLMs into one file serving OpenAI, Anthropic and Ollama APIs — p3rdy · 2026-10-07
- FastH3 V2 generates 5s video with audio in 15s on a single RTX 5090 — NVIDIAAI · 2026-10-07