GLM 5.3 Flash launches on Phala with GPU TEE private inference

bgmshana · x · 2026-08-29

Z.ai's GLM 5.3 Flash model is now available on the Phala network. It is a 320B MoE model with 18B active parameters, designed for efficient coding, long-horizon agent tasks, visual understanding, and long-context inference.

Deployment Features:

Pricing:

Original post →

More from Infra

Infra channel →