DeepSeek-v4-Pro Generates 130M Tokens for $1, 77x Cheaper Than Opus
bookwormengr · x · 2026-08-27
The Interact-vLLM team utilized Prefill-Decode disaggregation to generate 130 million tokens using DeepSeek-v4-Pro for just $1 at production scale (2000 GPUs). This is 77 times cheaper than Opus with similar cache hit rates.
More from Infra
- Nvidia guides $108B Q3 revenue, doubling growth even without China data center sales — inductionheads · 2026-08-27
- Palantir Karp: Serious enterprises need to own their models and infrastructure — JosephJacks_ · 2026-08-27
- Economics of Becoming an OpenRouter Provider with H200 Nodes — ell-hol1 · 2026-08-27
- Analysis layers exist between Nvidia sales and enterprise ROI — iamKierraD · 2026-08-27
- Apple still leads in laptop processor performance years after M1, leaving Intel and AMD behind — lemire · 2026-08-27
- Custom vLLM INT8 stack hits 972 tok/s on Qwen 27B with 4x MI100 ($6.5k rig) — 1ncehost · 2026-08-27