Why DeepSeek's API Is So Cheap: Tiny Model Size Boosts Single-Chip Throughput
AravSrinivas · x · 2026-08-03
Addressing questions about DeepSeek's extremely low API pricing, the author points out that their core advantage lies in a ridiculously small model size. Compared to Anthropic's Opus and Sonnet, DeepSeek's model is 5x to 10x smaller, allowing what used to require an 8-chip node to run on a single chip.
Because it's so small, the model runs extremely fast and can multi-serve multiple users on a single chip while maintaining decent speeds. By the author's math, DeepSeek can handle roughly 40x more traffic than Opus using the exact same compute. This high inference efficiency is why they can stay profitable even at crazy low prices.
Related event: DeepSeek's Low-Priced API Profitability Driven by Ultra-Small Model(2 posts)→
More from Infra
- Engineer's Reminder: Serve Models at Their Original Training Precision — andrew_n_carr · 2026-08-03
- Local Deployment Deep-Dive: Impact of KV Cache Precision on DeepSeek Models — esw123 · 2026-08-03
- OpenAI Said to Discuss $250B Nvidia Backstop for 10GW Data Center — Beth_Kindig · 2026-08-03
- Compute Squeeze May Force Neo-Labs to Open-Source Frontier Models — gorkem · 2026-08-03
- Persisting KV Cache on Free ARM: 15x Faster LLM Prefill — Annual_Manner_5901 · 2026-08-03
- TensorSharp Benchmark: Speculative Decoding Doubles DeepSeek Speed — fuzhongkai · 2026-08-03