Together AI launches serverless inference for Qwen3.8-2.4T-A95B with 99.9% SLA

togethercompute · x · 2026-08-13

Together AI's Serverless Inference now supports Qwen3.8-2.4T-A95B, offering managed high-throughput for coding and agentic workloads, with 256K context, 128K output, adjustable reasoning effort, and 99.9% SLA.

Related event: Together AI Launches Serverless Inference for Qwen3.8-2.4T-A95B(2 posts)→

Original post →

More from Infra

Infra channel →