DeepSeek-V4-Flash tops OpenRouter, reshaping LLM pricing with extreme cost-efficiency

创业邦 · wechat · 2026-08-03

Following the release of DeepSeek-V4-Flash, its extreme cost-efficiency has rapidly shifted the usage patterns of global developers. Recent data shows it has topped the API usage on OpenRouter, with total calls to Chinese models surpassing US closed-source models for consecutive weeks, prompting many overseas developers to migrate their production environments.

The article introduces the concept of a "Kill Line," representing the critical threshold balancing performance and price. Through underlying engineering optimizations (e.g., efficient inference architecture, dynamic KV caching), DeepSeek has drastically reduced per-token costs. Its output pricing is only a fraction of competitors like GPT-5.6 Luna, suffocating products that are weaker in performance yet higher in cost, and forcing rivals like OpenAI and Mistral to cut prices.

The core metric of LLM competition is shifting from raw parameters and benchmarks to "effective output per unit cost." This global market reshuffle triggered by cost-efficiency is far from over, marking the beginning of a long-term battle over model iteration, global service, and sustainable commercialization.

Related event: DeepSeek-V4-Flash Tops OpenRouter, Reshaping Model Pricing(2 posts)→

Original post →

More from Venture

Venture channel →