DeepSeek v4.1 flash TTFT comparison: Together AI crushes rivals on pre-warmed queries
zhyncs42 · x · 2026-09-18
A developer shares time-to-first-token comparisons for DeepSeek v4.1 flash across inference providers, saying Together AI crushes everyone on pre-warmed queries. Full numbers are in the attached chart.
More from Infra
- 600 tok/s single-request on Qwen 35B with Ninfer on an RTX Pro 6000 — CharlesStross · 2026-09-18
- Cadence sees India's EDA market doubling to $7.82B by 2031 — bookwormengr · 2026-09-18
- Google Open-Sources Agent Substrate on GKE: 10x Density, 1,000+ Dormant Agents per Host — blaizedsouza · 2026-09-18
- Redditor crams six V100 GPUs into a standard full-tower case for local LLM inference — Odd_Caterpillar_2994 · 2026-09-18
- Crusoe raises $3.9B at $30.9B valuation to build data centers and modular AI factories — TechCrunch AI · 2026-09-18
- A 2.5-hour first-principles primer on the semiconductor supply chain worth your time — blaizedsouza · 2026-09-18