Cohere moves to NVIDIA Blackwell, cutting token costs and TTFT by 30–50%
cohere · x · 2026-09-10
- NVIDIA released an "AI Tokenomics" white paper framing inference economics in four pillars: token utility, demand forecasting, supply optimization, and monetization, with case studies on Cohere, Perplexity, and Canva plus a demand-forecasting worksheet.
- Core thesis: data centers are becoming "token factories"; inference is now the dominant workload, and enterprises must learn to produce, manage, and monetize tokens.
- Cohere's case study claims that moving production workloads to NVIDIA Blackwell lowered token costs, increased throughput, and cut time-to-first-token by 30–50% on many workloads.
More from Infra
- Adaptive VRAM governor keeps 130M model training for 1M steps on an 8GB GPU without OOM — uBazzyZ- · 2026-09-10
- Epoch: Top AI firms' compute grows 4x/year; OpenAI up nearly 20x since 2023 — Jsevillamol · 2026-09-10
- Qwen3.8 Flash on 128GB Strix Halo: memory math and perplexity of Q4+Q8 n-gram hybrid quants — MarkoMarjamaa · 2026-09-10
- Agent auto-builds a custom Dockerfile to optimize new workspace start times — lucasmeijer · 2026-09-10
- OpenAI compute has grown ~20x since 2023; leading labs scale 4x/year — The Verge AI · 2026-09-10
- Keras ships ZeroModels: 100+ model families in pure Keras 3, runnable on any backend — fchollet · 2026-09-10