The Debate Over DeepSeek-Level Pricing and the Grok Serving Stack
teortaxesTex · x · 2026-07-07
Addressing claims about offering inference at "DeepSeek-level killer prices," teortaxesTex pushes back: assertions that xAI runs Grok using SGLang or that third parties claim costs are 5x lower than DeepSeek's own API are exaggerated. DeepSeek has been utilizing these serving optimization techniques since the V3 era, and the Grok API is actually more expensive. The discussion touches on inference serving stack optimizations and the economics of compute.
More from Infra
- llama.cpp lands Flash Attention tuning for RDNA4, big prefill gains on AMD — pmttyji · 2026-09-11
- Your p99 latency benchmark may be lying: a deep dive into coordinated omission — Franc0Fernand0 · 2026-09-11
- Running MiniMax H3 on 12GB VRAM: quantization, Turbo LoRAs and attention backends compared — Possible_Mood676 · 2026-09-11
- Spomin: live KV cache compaction squeezes 500k tokens of context into 180k resident — wgaca2 · 2026-09-11
- PiPNN nearest-neighbor search wins three awards, up to 78x faster index building — khademinori · 2026-09-11
- M.2-Oculink eGPU Link Silently Downgrades to PCIe Gen1 — Here's How to Check — El_90 · 2026-09-11