DeepSeek-V4-Flash Specs and Benchmarks Revealed: 284B MoE, 1M Token Context

ollama · x · 2026-08-05

Ollama's model page details the specs and benchmarks for DeepSeek-V4-Flash. The model features a Mixture-of-Experts architecture with 284B total and 13B activated parameters, supporting a massive 1 million token context window.

It offers three thinking modes: No thinking, Thinking, and Max thinking. Benchmarks show strong performance in its Max mode on tests like LiveCodeBench and HLE, with the V4-Pro version pushing boundaries even further.

Related event: DeepSeek-V4-Flash Tops Ollama Growth Chart(2 posts)→

Original post →

More from Models

Models channel →