GLM 5.3 Flash launches on Phala with GPU TEE private inference
bgmshana · x · 2026-08-29
Z.ai's GLM 5.3 Flash model is now available on the Phala network. It is a 320B MoE model with 18B active parameters, designed for efficient coding, long-horizon agent tasks, visual understanding, and long-context inference.
Deployment Features:
- Served via a TDX-attested GPU TEE (Trusted Execution Environment) for private and verifiable inference.
- Supports tool calling and JSON mode.
Pricing:
- Input: $0.15 / 1M tokens
- Output: $0.50 / 1M tokens
- Context Window: 1M tokens
More from Infra
- Ollama's New Claude Toggle Lets You Run Local Models Inside Claude Desktop — dr_cintas · 2026-08-29
- AI datacenters face visceral physical backlash as expansion meets local resistance — AccBalanced · 2026-08-29
- Running Qwen 3.8 MoE on a single DGX Spark: A practical recipe — QuixiAI · 2026-08-29
- Qwen 3.8 Flash NVFP4 deployment config tested on single DGX — QuixiAI · 2026-08-29
- AMD releases ROCm 10.0, built for the Age of Agentic AI — pmttyji · 2026-08-29
- Kyndryl and Broadcom bet on private AI clouds, with certified talent as the key — DavidLinthicum · 2026-08-29