DeepSeek-V4-Flash Specs and Benchmarks Revealed: 284B MoE, 1M Token Context
ollama · x · 2026-08-05
Ollama's model page details the specs and benchmarks for DeepSeek-V4-Flash. The model features a Mixture-of-Experts architecture with 284B total and 13B activated parameters, supporting a massive 1 million token context window.
It offers three thinking modes: No thinking, Thinking, and Max thinking. Benchmarks show strong performance in its Max mode on tests like LiveCodeBench and HLE, with the V4-Pro version pushing boundaries even further.
Related event: DeepSeek-V4-Flash Tops Ollama Growth Chart(2 posts)→
More from Models
- Claude Cites Expert Credentials to Bypass Its Own Safety Gate, Gets Blocked Anyway — matthew_d_green · 2026-08-05
- DeepSeek-V4-Flash Tested: $0.31 for Tasks That Cost $35 on Other Models — Teknium · 2026-08-05
- Anthropic and OpenAI Internally Months Ahead of Public Models — haider1 · 2026-08-05
- DeepSeek V4 Flash Offered at 90% Off on Vercel, Touted as Opus 4 Rival — cramforce · 2026-08-05
- OpenAI Testing Dedicated Download Page for Life Sciences Model GPT-Rosalind Codex — testingcatalog · 2026-08-05
- DeepSeek-V4-Flash Becomes Fastest Growing Model on Ollama with Zero Data Retention — ollama · 2026-08-05