Project Maya runs GLM-5.3-Flash (321B MoE) locally at up to 118 tok/s on 4×4090s
inthesearchof · reddit · 2026-10-12
Project Maya (open source, 365 GitHub stars) runs Zai's GLM-5.3-Flash — a 321B-parameter MoE with 18B active and 1M context, MIT-licensed — on your own GPUs, built on Strata.
- Tiered offloading: hot experts on GPU, next tier in RAM, rest on NVMe SSD, shuffled as you chat
- Auto-tuning: benchmarks your GPU/CPU/RAM/SSD and configures itself; 1–16 NVIDIA GPUs, experimental AMD/Windows
- Speed: up to 118 tok/s on 4×RTX 4090s, 30+ tok/s on a single RTX 5090
- Quality: custom quants retain 97.7–99.2% of the FP8 model's zero-shot accuracy
- Integration: OpenAI/Anthropic-compatible API with tool calls, ready for coding agents; bundled browser chat and monitor
The poster reports going from 10 to 30 tok/s and finds low-quant GLM-5.3 more enjoyable than Qwen 3.8 Flash Next so far.
More from Infra
- Is local AI trending toward GPU-interconnect-friendly workloads? — Dathide · 2026-10-12
- What's the best local coding setup for 16GB VRAM right now? — ECrispy · 2026-10-12
- Usage dashboard shows 1,100+ cloud VMs spun up, one account with 738 machines — aniketmaurya · 2026-10-12
- Local AI on AMD Strix Halo Writes Full Tech Specs: 5-10x Slower but It Works — julianharris · 2026-10-12
- jax-graft: an AI-built JAX backend runs JAX on Apple Silicon GPUs — twiecki · 2026-10-12
- GamePause: open-source tray app auto-unloads local LLMs when gaming, frees 17.4GB VRAM — zainfear · 2026-10-12