GLM 5.3 Flash Served on 2 DGX Sparks: Open Recipe Hits 77.6 tok/s with 3
EAccelerate_42 · x · 2026-10-08
Mia's AI Lab released v1.10 of its open recipe to serve GLM 5.3 Flash locally on two (or three, experimental) NVIDIA DGX Sparks via TensorFold.
Highlights
- Custom EXL3 4-bit quant, TensorFold v0.6.0 + 96 patches
- Works directly with Claude Code, OpenAI-compatible and Anthropic SDKs, with tool calling
- Much faster long chats (500k+ tokens); opt-in disk cache resumes long chats in 0.2s instead of 11s, even after restart
- Sampling now as fast as greedy decoding; FP8 KV cache, RoCE interconnect, vision input, full 1M-token context, Apache 2.0
- A third Spark pushes throughput to 77.6 tok/s
One command sets up the cluster and starts the server; the free recipe and agent-install prompt are public. Users report it runs well in the DeepSeek harness Mac app.
More from Infra
- Running an LLM town with 800+ persistent agents: concurrency, caching, and costs — Low_Bad_6585 · 2026-10-08
- AI borrowing costs hit ~11% as JPMorgan markets $5B Volta loan for 36,000 Nvidia GPUs — mjdramstead · 2026-10-08
- tinygrad launches new tinybox deep learning rig, configurable up to 4 GPUs — GiorgioPatrini · 2026-10-08
- Inference Overtakes Training as Biggest AI Market, Reshaping Data Center Buildouts — FinanceYF5 · 2026-10-08
- New llama.cpp PR assigns four GDN state columns per warp for another Qwen 3.x prefill speedup — jacek2023 · 2026-10-08
- MIT Tech Review: building a safer path to autonomous industrial AI — nordicinst · 2026-10-08