DeepSeek API adds v4-pro and v4-flash with 1M context and legacy endpoint retirement
teortaxesTex · x · 2026-07-22
DeepSeek’s API now supports deepseek-v4-pro and deepseek-v4-flash while keeping the same base URL. The models expose both OpenAI ChatCompletions and Anthropic-compatible APIs, and each supports 1M context plus two modes: Thinking and Non-Thinking.
The quote also highlights the planned retirement of legacy deepseek-chat and deepseek-reasoner endpoints on Jul 24, 2026, 15:59 UTC. Pricing shown in the screenshot puts v4-pro at $0.145 / $1.74 / $3.48 for input cache hit, input cache miss, and output respectively, while v4-flash is much cheaper at $0.028 / $0.14 / $0.28.
More from Infra
- Cloudflare-style infra is making agent-first apps feel radically easier to build — threepointone · 2026-07-22
- OpenFPM CUDA-style kernels now run on Apple Silicon GPUs via Metal — Scobleizer · 2026-07-22
- SkewAdam cuts MoE optimizer memory by 97.4% and fits 6.78B on a 40GB GPU — Kooky-Ad-4124 · 2026-07-22
- Samsung is said to weigh a €1B Mistral investment at a €20B valuation — rohanpaul_ai · 2026-07-22
- NVIDIA and ETH Zürich cut small-message AllReduce latency by deleting barriers — thoefler · 2026-07-22
- Grok Build adds token usage, batching and diagnostics for developers — elonmusk · 2026-07-22