The Debate Over DeepSeek-Level Pricing and the Grok Serving Stack
teortaxesTex · x · 2026-07-07
Addressing claims about offering inference at "DeepSeek-level killer prices," teortaxesTex pushes back: assertions that xAI runs Grok using SGLang or that third parties claim costs are 5x lower than DeepSeek's own API are exaggerated. DeepSeek has been utilizing these serving optimization techniques since the V3 era, and the Grok API is actually more expensive. The discussion touches on inference serving stack optimizations and the economics of compute.
More from Infra
- YC talk on BCI x AI says infrastructure is what really determines speed — garrytan · 2026-07-27
- A 13B model ran on a no-GPU PC by paging weights from SSD via llama.cpp — ID_R_McGregor · 2026-07-27
- llama.cpp warns that GGUFs made before a recent change must be regenerated — EconomySerious · 2026-07-27
- RTX 5090 local tests show Qwen Q6 can drop to 15 tok/s at 80k context — LFAdvice7984 · 2026-07-27
- Surprising Ubuntu Setup: NVIDIA 5090 PC Becomes the Easiest AI Rig — _xjdr · 2026-07-27
- TSMC reportedly plans 5%–10% price hikes in 2027 to cover rising costs — Beth_Kindig · 2026-07-27