DeepSeek-V4-Flash Rolls Out on Ollama Cloud with 120+ Output TPS

ollama · x · 2026-08-08

Ollama has fully rolled out DeepSeek-V4-Flash-0731 as the new default for its cloud platform. The model combines efficiency and frontier-level performance, achieving 120+ output tokens per second on Ollama's cloud. It also features zero data retention hosting in the US and Europe, supporting long-running sessions with popular coding harnesses.

Related event: DeepSeek-V4-Flash Released Across Major Platforms(4 posts)→

Original post →

More from Models

Models channel →