DeepSeek-V4-Flash Rolls Out on Ollama Cloud with 120+ Output TPS
ollama · x · 2026-08-08
Ollama has fully rolled out DeepSeek-V4-Flash-0731 as the new default for its cloud platform. The model combines efficiency and frontier-level performance, achieving 120+ output tokens per second on Ollama's cloud. It also features zero data retention hosting in the US and Europe, supporting long-running sessions with popular coding harnesses.
Related event: DeepSeek-V4-Flash Released Across Major Platforms(4 posts)→
More from Models
- Questioning AI Reasoning: Are Baseline Capabilities Still Climbing With Reasoning Turned Off? — inductionheads · 2026-08-08
- GPT-5.6 Sol Praised for Speeding Up Cybersecurity Incident Response — gdb · 2026-08-08
- OpenAI and ElevenLabs Adopt Google's SynthID Audio Watermarking Technology — pushmeet · 2026-08-08
- DeepSeek V4 Lacks Vision: Model Builds Brightness Grid to Self-Correct — teortaxesTex · 2026-08-08
- MiniMax H3 Training Hits a Wall: Community Reports Crashes on Large Datasets — Sorry_Warthog_4910 · 2026-08-08
- Leaked Claude 4 Test Shows 384k Max Output Tokens, Hints Higher Limits — teortaxesTex · 2026-08-08