GLM 5.3 Flash Served on 2 DGX Sparks: Open Recipe Hits 77.6 tok/s with 3

EAccelerate_42 · x · 2026-10-08

Mia's AI Lab released v1.10 of its open recipe to serve GLM 5.3 Flash locally on two (or three, experimental) NVIDIA DGX Sparks via TensorFold.

Highlights

One command sets up the cluster and starts the server; the free recipe and agent-install prompt are public. Users report it runs well in the DeepSeek harness Mac app.

Original post →

More from Infra

Infra channel →