GLM 5.3 Flash goes live on W&B serverless inference: 1M context, vision, $0.50/M output
wandb · x · 2026-09-09
A developer reports GLM 5.3 Flash is now live on W&B serverless inference, backed by CoreWeave, offering 1M token context, vision support, $0.50 per million output tokens, and fast speeds. The same author also spotted DeepSeek V4 live on the platform earlier. Third-party availability change for open-weight models; details unconfirmed by vendors.
More from Infra
- VeloxML: open-source engine self-hosts LLMs in your own cloud with one command and scale-to-zero — pm19191 · 2026-09-09
- Virginia Hosts 13% of World's Data Centers Yet Has Cheaper Power Than Average — Scobleizer · 2026-09-09
- NVIDIA's Online Draft Co-Training Speeds Speculative Decoding in Long-Context RL Post-Training — nvidia · 2026-09-09
- PostgreSQL 19 ships property graphs: query relationships on existing tables without joins — arpit_bhayani · 2026-09-09
- Why AI Hasn't Boosted Growth Yet: It's a Function of Global Inference Capacity — zephyr_z9 · 2026-09-09
- PyTorch highlights cross-community collaboration at KubeCon + PyTorchCon China 2026 — PyTorch · 2026-09-09