DeepSeek-V4-Flash Hits Ollama Cloud with 120+ Output Tokens/sec

zhyncs42 · x · 2026-08-11

Ollama has fully rolled out DeepSeek-V4-Flash on its cloud platform. The model combines speed, efficiency, and frontier-level performance, achieving 120+ output tokens/sec on Ollama's cloud.

It features open weights and zero data retention hosting in the US and Europe. Pro and Max plan users get generous usage allowances to support long-running, uninterrupted sessions for coding harnesses.

Original post →

More from Models

Models channel →