Together AI launches serverless inference for Qwen3.8-2.4T-A95B with 99.9% SLA
togethercompute · x · 2026-08-13
Together AI's Serverless Inference now supports Qwen3.8-2.4T-A95B, offering managed high-throughput for coding and agentic workloads, with 256K context, 128K output, adjustable reasoning effort, and 99.9% SLA.
Related event: Together AI Launches Serverless Inference for Qwen3.8-2.4T-A95B(2 posts)→
More from Infra
- Running DeepSeek V4 Flash Locally on 2x DGX Sparks Delivers Prosumer-Grade Performance — andrewchen · 2026-08-13
- Is Local Generative AI Worth It Anymore? Developers Struggle Against Closed Cloud Models — ImaginaryEffective63 · 2026-08-13
- Intel Razor Lake AX Info Surfaces, Targeting AMD's Future Local AI Chips — Terminator857 · 2026-08-13
- Investor Burry Shorts Compute Stocks, Sparking Debate Over AI Compute Shortage — inductionheads · 2026-08-13
- $14.6B AI Compute Bet: Jane Street Needs 20.3% Annual Yield to Break Even — adrianscottcom · 2026-08-13
- AI's Insatiable Appetite: Someone Desperately Seeking a 2PB WEKA Server — charles_irl · 2026-08-13