DeepSeek's Tiny Model Size Yields 40x Opus Throughput, Enabling Low Prices
gaganghotra_ · x · 2026-08-02
Dispelling rumors of market-dumping, Jun Song explains that DeepSeek's highly profitable, low-cost API is driven by its ridiculously small model size relative to performance.
- Size Advantage: The model is 10x smaller than Claude Opus and 5x smaller than Sonnet while maintaining strong capabilities.
- Compute Efficiency: Workloads that previously required an 8-chip node can now run on a single chip at high speeds.
- Massive Throughput: By multi-serving users on a single chip efficiently, it can handle roughly 40x more traffic than Opus using the exact same compute footprint.
This extreme size-to-performance ratio is the core reason they can maintain profitability at aggressively low prices.
More from Infra
- DeepSeek V4 Flash Prefill Speed Boost: Downgrade to CUDA 13.1 — fragment_me · 2026-08-03
- Cornell Releases Roadmap for Parallel Programming and HPC Concepts — thehiphopswami · 2026-08-03
- Running DeepSeek V4 Flash 155G on DGX Spark: 2-bit Quantization & MTP Benchmarks — Puzzleheaded_Base302 · 2026-08-03
- DFlash: Parallel Speculative Decoding via Lightweight Block Diffusion — cneuralnetwork · 2026-08-03
- Mainstream LLM API Uptime Tested: Anthropic Lags, Kimi K3 Hits 99.4% — weswinder · 2026-08-02
- Local LLM on Mac: M2 Ultra 192GB Long-Context Inference Benchmarks — Badger-Purple · 2026-08-02