DeepSeek-v4-Pro Generates 130M Tokens for $1, 77x Cheaper Than Opus

bookwormengr · x · 2026-08-27

The Interact-vLLM team utilized Prefill-Decode disaggregation to generate 130 million tokens using DeepSeek-v4-Pro for just $1 at production scale (2000 GPUs). This is 77 times cheaper than Opus with similar cache hit rates.

Original post →

More from Infra

Infra channel →