Baseten ships GLM-5.3 Fast: speed-optimized open-weight model for real-time workloads
baseten · x · 2026-09-03
Baseten released GLM-5.3 Fast, a speed-optimized version of Z AI's GLM-5.3 aimed at real-time workloads needing consistent performance at higher TPS.
- Model: 753B-A40B MoE, same 744B-A40B base as GLM-5.2; all gains come from scaled post-training on diverse, realistic task environments (full codebases, docs, testing tools, multi-step workflows)
- Capabilities: Z AI's strongest coding model yet — Terminal-Bench 3.0 jumps from 4.6% to 28.3% over GLM-5.2, plus notable gains in vulnerability discovery and security analysis
- Pricing: $2.10/M input tokens, $0.21/M cached input, $6.60/M output
- Access: OpenAI-compatible API on inference.baseten.co, with three thinking effort levels (max recommended for coding)
More from Models
- IBM releases Granite 4.2: free open-source models built for AI agents, runs locally — krvarshney · 2026-09-03
- Astra's rumored looped transformer gets a technical debunk, with Oriol Vinyals citing Universal Transformer — OriolVinyalsML · 2026-09-03
- Muse model now testable in opencode, Cursor support still uncertain — talkaboutdesign · 2026-09-03
- Grok 4.7 reportedly lands in 10 days: ~2.1T params, 40% scale jump, claims to top all models — tetsuoai · 2026-09-03
- Users Report Claude Racking Up Daily Mistakes and Hallucinations — lilyraynyc · 2026-09-03
- First run of Gemini 3.8 Flash fails: model keeps thinking until it times out — rickasaurus · 2026-09-03