Reuters: DeepSeek Developing Own AI Inference Chip
机器之心 · wechat · 2026-07-08
Reuters reported that DeepSeek is developing its own AI chip designed specifically for the inference phase to reduce its reliance on Nvidia and domestic chips. Following the news, Nvidia fell about 1.6% in pre-market trading. This is a major strategic shift for DeepSeek. Currently in its early stages, the company is in talks with chip design, wafer foundry, and storage manufacturers, and is privately ramping up the recruitment of chip design engineers.
DeepSeek previously used Nvidia H800 to train the R1 base model, switched to Huawei Ascend after export controls at the end of 2023, and the V4 released in April has already been adapted for Ascend. Self-developed chips have become a trend among top model vendors: OpenAI released its first custom inference chip in June (designed by Broadcom, manufactured by TSMC), and Anthropic was also reported to have started self-development and engaged with Samsung.
DeepSeek's self-developed chip still faces multiple challenges: moving from design to mass production requires years and massive capital, obtaining advanced processes and HBM is difficult due to export controls, and it must compete against the Nvidia CUDA ecosystem. This move coincides with its first introduction of external capital, planning to raise $7 billion in its initial funding round at a valuation of $52-59 billion.
More from Infra
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- 80% of the DIY LLM inference hype posters have already quit — it's brutally hard systems work — abhijithneil · 2026-09-11
- Hugging Face's Ultra Scale Playbook: a free book on training LLMs on GPU clusters — mdancho84 · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11