GLM-5.3 live on Phala inside TDX-attested GPU TEE, 1M context at $1.40/M input
bgmshana · x · 2026-09-01
Phala Network announced that Z.ai's GLM-5.3 is now served on its platform inside a TDX-attested GPU TEE for private inference. The model targets complex software engineering and long-horizon agent tasks, with a 1M-token context window, tool use, and structured outputs. Pricing: $1.40/M input, $4.40/M output. Phala reports 1.2s TTFT, 56 tps throughput, and 97.81% uptime over the last 72 hours. The catalog also includes GLM-5.3 Flash (320B MoE, $0.15/M input) and other TEE-deployed models.
More from Models
- Z.ai Releases GLM-5.3-Flash: 320B Params, 1M Context, and NVFP4 Quantization — alejandroll10 · 2026-09-01
- Open Source Models Shift to Revenue Sharing and Licensing — zephyr_z9 · 2026-09-01
- Has anyone tuned a model to operate exclusively in E-prime yet? — cephaloform · 2026-09-01
- Heavy users report Claude quality dropping over the past week: eager to execute, no more clarifying questions — Numerous_Leopard_522 · 2026-09-01
- User seeks best LLM for CLI coding on single 3080 Ti — -samae1- · 2026-09-01
- Lan Hackathon Review: Qwen Excels in Physics/Engineering, K3 in General Intelligence — 葬AI · 2026-09-01