Local Deployment of GLM-5.3-Flash: 206 tok/s and 1M Context on DGX Station

funding__secured · reddit · 2026-08-27

The author successfully deployed and tested the GLM-5.3-Flash model on a DGX Station GB300.

Performance:

Deployment Configuration (Docker):

A complete docker run command is provided with key parameter optimizations:

Gotcha:

The current image has a bug preventing automatic model downloads. Weights must be pre-downloaded locally and pointed to via the --model flag.

Related event: GLM-5.3 Flash tests show near-GPT-5.6 performance at rock-bottom cost(6 posts)→

Original post →

More from Infra

Infra channel →