Alibaba's Qwen3.8-Max Released, TokenSpeed Announces Inference Optimization
zhyncs42 · x · 2026-08-03
Alibaba's Qwen team has officially released Qwen3.8-Max, claiming it sets a new bar for coding and cowork.
Simultaneously, inference provider TokenSpeed announced Day-0 support for the model. They are currently optimizing multi-node inference to improve per-GPU TPM at high per-user TPS.
Related event: Alibaba Releases 2.4T Parameter Qwen3.8-Max, Open Weights Next Week(13 posts)→
More from Infra
- Google TPU Demand Surge: Projected to Reach 15M Units by 2028 — SumitGup · 2026-08-03
- Run 2.78T Parameter Kimi K3 on a Single CPU in 8.24GB RAM — Saboo_Shubham_ · 2026-08-03
- tinygrad Teases Local Deployment Product, Hints at Upcoming Qwen3.6-27B — max_paperclips · 2026-08-03
- AMD MI355X Beats NVIDIA B200 in Kimi K3 Deployment with 952 tok/s — adrianscottcom · 2026-08-03
- Qwen3.8-27B Open Weights Coming, Runs Locally on 17GB RAM — danielhanchen · 2026-08-03
- AirLLM Breaks VRAM Barrier: Runs 70B LLMs on a Single 4GB GPU — techNmak · 2026-08-03