GLM 5.3 Flash goes live on W&B serverless inference: 1M context, vision, $0.50/M output

wandb · x · 2026-09-09

A developer reports GLM 5.3 Flash is now live on W&B serverless inference, backed by CoreWeave, offering 1M token context, vision support, $0.50 per million output tokens, and fast speeds. The same author also spotted DeepSeek V4 live on the platform earlier. Third-party availability change for open-weight models; details unconfirmed by vendors.

Original post →

More from Infra

Infra channel →