DeepSeek Beats OpenAI in Inference Margins via SOTA KV Cache Optimization

basedjensen · x · 2026-08-01

Despite OpenAI's Luna model being significantly larger, DeepSeek achieves higher inference margins on V4 Flash compared to OpenAI using Blackwells for Luna 5.6. This advantage is attributed to DeepSeek's state-of-the-art KV cache offload system and highly optimized kernel engineering.

Related event: DeepSeek on Huawei Ascend Beats OpenAI in Inference Profitability(2 posts)→

Original post →

More from Infra

Infra channel →