Baidu serves DeepSeek V4 Flash at crazy fast speeds and low prices
NielsRogge · x · 2026-08-19
Users have observed that Baidu is serving the DeepSeek V4 Flash 0731 model at extremely fast speeds and very low prices. The model is a sparse mixture-of-experts with 13B active parameters out of 284B total, designed for coding, reasoning, and agent workflows. OpenRouter lists the price at $0.0786 per 1M input tokens with 1M context support.
More from Infra
- Dev mocks vague NVIDIA PTX docs, credits Modular for coining 'Core Matrix' — nrehiew_ · 2026-08-19
- DFlash 2 available for Qwen 3.8 27B and Muse Glimmer — rerri · 2026-08-19
- AI hyperscalers' $308B debt buildout is pushing up Treasury yields — tszzl · 2026-08-19
- Deep Dive into Train-Infer Mismatches: Causes from Floating-Point to Architecture — nrehiew_ · 2026-08-19
- llama.cpp reasoning-preserve flag risks context bloat — anderspitman · 2026-08-19
- TradingView MCP Server: Real-time market data for Claude & ChatGPT — tom_doerr · 2026-08-19