Project Maya runs GLM-5.3-Flash (321B MoE) locally at up to 118 tok/s on 4×4090s

inthesearchof · reddit · 2026-10-12

Project Maya (open source, 365 GitHub stars) runs Zai's GLM-5.3-Flash — a 321B-parameter MoE with 18B active and 1M context, MIT-licensed — on your own GPUs, built on Strata.

The poster reports going from 10 to 30 tok/s and finds low-quant GLM-5.3 more enjoyable than Qwen 3.8 Flash Next so far.

Original post →

More from Infra

Infra channel →