FULL STORY

DeepSeek V4.1 Flash: From Limited Beta to Open-Source Release

After a limited beta on Sept 8, DeepSeek officially released and open-sourced V4.1 Flash on Sept 10 with price cuts. Benchmarks show near-frontier performance, a 4x smaller KV cache, and the top spot on the Vals open-weight leaderboard.

2026-09-08 ~ 2026-09-11 · 15 episodes · 154 posts

Episode 1 · DeepSeek V4.1 Flash Opens Limited Internal Beta with New Architecture and Native Multimodality (2026-09-08, 18 posts)

Starting September 8, multiple sources reported that DeepSeek quietly opened a limited internal beta for V4.1 Flash, a "middle version" with model ID deepseek-v4.1-flash-expires-on-0910, the slug indicating expiry on September 10. Key upgrades include a brand-new architecture and native multimodality, with the official line claiming stronger capability, faster speed, and lower cost. Calling conventions, billing, and concurrency rules are now largely clear, developer benchmarks have emerged, and the model has appeared on Hugging Face, though an official public announcement and full evaluations remain pending.

Confirmed

  • Jiqizhixin reported that DeepSeek officially announced the internal test of the V4.1 Flash middle version; APPSO, Zhidongxi and other outlets confirmed the same.
  • Usage: keep baseurl unchanged and set the model name to an ID beginning with deepseek-v4.1-flash-expires-on-09 (e.g., deepseek-v4.1-flash-expires-on-0910), where the date marks expiry.
  • According to vista8's relay of an official WeChat group notice, billing matches V4 Flash with max 20 concurrent requests per account; teortaxesTex's leaked info points to the same concurrency limit and a vision-capable build (V4-Flash-Vision [Intermediate]).
  • On pricing, MiaAIlab (relayed by MicahBerkley) described the price as extremely aggressive: roughly 58 million tokens per US dollar, available for about 2 days.
  • Third-party account LuminaBench spotted the intermediate build being tested via API; tokenbender suggested it precedes the official release.
  • On September 10, Kevin Kern tested the preview endpoint, measuring 300 tok/s and heavily parallelizing subagents; he noted the preview endpoint does not yet accept vision input. Zhidongxi reported output speeds up to 507 tokens/s.
  • On September 10, Reddit users found a deepseek-ai/DeepSeek-V4.1-Flash model page on Hugging Face, suggesting a quiet official upload; specs unconfirmed.

Unconfirmed

  • The news initially spread via bloggers such as op7418 and AGI Hunt from a DeepSeek WeChat group, unverified at the time; Jiqizhixin later framed it as officially announced, but a formal public announcement and full capability details are still pending.
  • @NFTChen (relayed by solyarisoftware) claimed V4.1 Flash would be officially released around September 10 Beijing time, with V4 Pro requests routed at Flash pricing; not confirmed by DeepSeek.
  • The preview endpoint's lack of vision input sits in tension with the "native multimodality" claim; whether the formal release adds vision remains to be seen.
  • teortaxesTex relayed community banter hoping it can at least beat GLM 5.3 Flash outside of vision—opinion only, no benchmark basis.

Why it matters

  • This is DeepSeek's first reported use of an "expires-on" middle-version limited beta with explicit billing and concurrency rules plus aggressive pricing, signaling a product cadence of rapid iteration through real developer workloads before formal release.
  • Native multimodality marks a major upgrade to DeepSeek's flagship line; with 300 tok/s measured, an appearance on Hugging Face, and the community benchmarking it against GLM 5.3 Flash, the formal release's multimodal capabilities deserve close attention.

Episode 2 · DeepSeek V4.1 Flash Tested: Blazing 350 Tokens/s but Still Experimental (2026-09-08, 2 posts)

Hands-on tests of DeepSeek's experimental V4.1-Flash model show decoding speeds of roughly 300-400 tokens per second at competitive prices, though reviewers note it remains early-stage and prone to overthinking.

Episode 3 · DeepSeek cuts V4-Flash API prices with new peak/off-peak billing from Sept 10 (2026-09-08, 6 posts)

DeepSeek officially announced a price cut for its V4-Flash (flash) API series effective Sept 10, 12:00 Beijing time, restoring input prices to pre-hike levels and switching to peak/off-peak billing, with the largest discounts in off-peak hours. The new scheme replaces the expiring "Intermediate" V4.1 pricing; the price sheet had leaked via community channels a day earlier.

Confirmed

  • DeepSeek platform announcement: V4-Flash price cut from Sept 10, 12:00 (reported by 机器之心, 量子位, aigclink).
  • Off-peak pricing (yuan per million tokens): cache-hit input 0.02 (down 60% from 0.05); cache-miss input 1 (down 33.33% from 1.5); output 4 (down 11.11% from 4.5).
  • Peak hours are weekdays 9:00–12:00 and 14:00–18:00; the rest is off-peak. Peak prices are higher, per the official schedule.
  • The new scheme replaces the expiring "Intermediate" V4.1 pricing (per teortaxesTex).
  • michaelsoftbinbows noted that DeepSeek V4 Flash (0731) via openinference on OpenRouter costs only $0.05 input / $0.16 output per 1M tokens off-peak, below official pricing (from $0.22), drawing wide discussion.

Unconfirmed

  • WeChat account 赛博禅心 mentioned a 20-concurrency limit accompanying the new scheme, flagged as pending official confirmation; the official announcement is silent on this.

Timeline

  • Sept 8: teortaxesTex summarized the leaked pricing; 赛博禅心 posted the cut with effect time and concurrency details, unconfirmed.
  • Sept 9: Official announcement confirmed via 量子位, 机器之心, aigclink; OpenRouter third-party low pricing sparked discussion.
  • Sept 10, 12:00: New prices take effect.

Why it matters

  • Input prices returning to pre-hike levels partially rolls back the earlier increase, directly benefiting developers and vendors relying on DeepSeek APIs.
  • Off-peak cache-hit pricing as low as 0.02 yuan per million tokens, combined with time-based billing, will push users toward off-peak calls and higher cache utilization, reflecting shifts in DeepSeek's capacity and scheduling.
  • The notable gap between official and third-party channels (e.g., OpenRouter) offers developers a cheaper alternative access path.

Episode 4 · DeepSeek V4.1 Flash tested across tasks: near-frontier performance at a fraction of the cost (2026-09-09, 15 posts)

DeepSeek V4.1 Flash, along with the V4.1 Flash Vision Beta released the same day, has been put through multiple rounds of independent testing in the days since launch. The consensus: its performance approaches or even matches frontier closed-source models in some areas, at just 1/100 to 1/30 of the cost.

Confirmed

  • Design benchmark: In a comparison of everyday design tasks, blogger NFTChen scored DeepSeek V4.1 Flash at 81.2—98% of GPT-6 Astra's 82.7 and ahead of Claude Fable 5.1's 80.3—at roughly 1.4% of Astra's cost; OpenDesign Arena leaderboard data shows similar value.
  • Cybersecurity benchmark: The pilvar222 team found V4.1 Flash rediscovers 65.6% of recent CVEs in the benchmark in a single run (vs. 55.2% for the old version), with pass@3 catching 84.4% (vs. 75%), matching frontier models.
  • Real-world bug fixing: Pawel Huryn tested with 2 repos and 105 hidden real bugs; V4.1 Flash landed on the Pareto frontier—for reference, Opus 5 (max) found 27 bugs at $51.33, while V4.1 Flash costs about 3.5% of Opus (1/28 the price). The author also added an FAQ addressing common objections.
  • Vision: On launch day, Reddit user cheezeerd compared V4.1 Flash Vision Beta with V4 Flash Vision across 5 vision tasks, with full video walkthroughs of the first draft and three revisions—concluding the new version produces noticeably better, more reliable output at lower API pricing.

Why it matters

  • Tests from multiple independent sources across different task types (design, security, code, vision) corroborate V4.1 Flash's value proposition, rather than resting on a single benchmark.
  • If the results replicate, ultra-cheap near-frontier models will further squeeze pricing for closed-source flagships and lower the barrier for developers and enterprises to switch.

Not yet confirmed

  • All results come from third-party individuals or teams with varying samples and methodologies; official metrics and large-scale replications have yet to appear.

Episode 5 · DeepSeek releases V4.1-Flash: 552B MoE beats flagships at low cost (2026-09-09, 56 posts)

On September 10, DeepSeek officially released and open-sourced its new model V4.1 Flash, rolling out simultaneously on web, app, and API, with the API model named deepseek-flash. This is one of the most notable new open-source model releases right now: performance reportedly surpasses their own flagship V4 Pro across the board, at a sharply lower price.

Confirmed

  • Architecture and scale: a 552B total-parameter asymmetric MoE model with a brand-new Causal-Encoder-Decoder architecture, activating only 8B parameters on the input side and 16B on the output side; supports up to 1 million tokens of context, with KV cache overhead cut by 7/8.
  • Benchmarks: per @ChrisGPT, V4.1 Flash beats GPT-5.6 Sol and Opus 5 on multiple agentic/coding benchmarks, e.g., CyberGym 88.1 vs 84.5; input pricing as low as $0.14/M.
  • V4 Pro retirement and routing: the official announcement states that since V4.1 Flash surpasses V4 Pro across performance, cost, speed, and total time, all V4 Pro requests will be automatically routed to V4.1 Flash and billed at the lower price until V4.1 Pro launches. Several users (e.g., @gaganghotra) have already received deprecation notices.
  • Older models sunset: the official API announcement discloses that legacy models like V4-Flash and V4-Flash-Vision-E will be retired.

Why it matters

  • V4.1 Flash matches and exceeds flagship V4 Pro performance with far fewer activated parameters and much lower cost—combined with its open-source release, this could further drive down industry inference prices and pressure closed-model vendors.
  • V4 Pro users should watch for behavior changes from the automatic routing; developers relying on its capabilities should evaluate V4.1 Flash's actual performance before switching.
  • The 1M-token context and 7/8 KV cache savings have direct implications for usability and cost in long-document and agentic workflow scenarios.

36 more related posts →

Episode 6 · DeepSeek V4.1 Flash Leak: 552B Asymmetric MoE Reportedly Rivals GPT-5.6 (2026-09-10, 20 posts)

On September 10, several bloggers collectively leaked what are claimed to be official benchmark results and technical details for DeepSeek V4.1 and V4.1 Flash; none of the information has been officially confirmed.

Confirmed

  • All information is unverified rumor, mainly from bloggers such as teortaxesTex, eliebakouch, and ChrisGPT, whose posts corroborate each other.
  • The model has 552B total parameters with a brand-new Causal-Encoder-Decoder architecture, 8B input activation and 16B output activation.
  • Leaked scores: TerminalBench 4.0 at 31.2, CyberGym 88.1 (vs. 84.5 for GPT-5.6 Sol), HLE 63.9; ChrisGPT claims it beats GPT-5.6 Sol on multiple agentic/coding benchmarks.
  • The technical report cited by eliebakouch shows V4.1 Flash is a natively vision-capable model trained on 45T tokens, with an architecture praised as the most novel in years.

Architecture and Engineering Highlights

  • KV cache compressed to 890 bytes/token, under 1GB for a million tokens; compared with V4-Flash, HBM requirements are down another 3.9x and SSD requirements down 8x, still unsurpassed half a year after release (per teortaxesTex).
  • According to teortaxesTex, V4.1 removes V4's stopgap patches: dropping HCA, generalizing CSA into a more universal primitive, eliminating the warmup stage of sparse attention, and adopting a simpler single-pass mHC.

Not Yet Confirmed

  • All scores, specs, and architecture details lack official sources; the final release may differ.

Why It Matters

  • If true, V4.1 directly challenges GPT-5.6 Sol on agentic/coding capability while slashing deployment memory requirements, with significant implications for inference costs in an open-source ecosystem.

Episode 7 · DeepSeek V4.1 Flash Benchmark Results Leak, Sparking Community Discussion (2026-09-10, 2 posts)

Early benchmark screenshots circulating on Reddit show DeepSeek V4.1 Flash achieving SOTA on the deepswe leaderboard alongside sol and fable5, with competitive terminal bench results, sparking community discussion.

Episode 8 · Bug Hunt Bench Tests 105 Real Bugs; DeepSeek V4.1 Flash Stands Out on Cost (2026-09-10, 8 posts)

Pawel Huryn (author of The Product Compass) launched and maintains the Bug Hunt Bench benchmark: 105 real bugs were deliberately planted in two production codebases, and frontier models including GPT-6, Claude, Grok, Gemini, DeepSeek, Kimi, and GLM were tasked with fixing them in their respective agentic configurations, ranked by blind review of the number of fixes, with only one prompt given per repository, a maximum score of 105, and results continuously updated on its GitHub leaderboard page.

Confirmed

  • GPT-6 Astra supplementary results: low tier 27/105, medium tier 34/105; these were previously released on September 4-5 and can be compared via filters on the leaderboard page.
  • DeepSeek-V4.1-Flash (high tier): fixed 24/105 at max effort, with a total cost of only $1.80 and a runtime of 42.6 minutes.

Why it matters

  • With its setup of pre-planted bugs in real repositories, blind review, and a single prompt, this benchmark is closer to real-world bug-fixing scenarios than traditional synthetic benchmarks.
  • Huryn highlighted that DeepSeek-V4.1-Flash fixed 24 bugs for roughly $1.80, delivering standout cost-effectiveness in its price class and providing valuable selection data for budget-conscious teams.

Episode 9 · Inside DeepSeek V4.1 Flash: YOCO at its core and KV cache reuse (2026-09-10, 7 posts)

On September 10, several bloggers focused on architectural changes in DeepSeek's new V4.1 model: DeepSeek is said to be returning to an encoder-decoder architecture, with multiple innovations at the architectural level and full open-sourcing—widely seen as continually raising the technical floor of the AI research ecosystem. Most of this information comes from paper details and KOL analysis; the official position is not yet unified, and the architecture's naming remains disputed.

Confirmed

  • DeepSeek has open-sourced the technologies related to the new architecture.
  • Astra Pro's deep dive into the V4.1 paper found that DeepSeek reuses KV caches across different checkpoints, prompting commenters to call it a "strange model factory."
  • DeepSeek cites the 2024 YoCo (You Only Cache Once) paper as the main inspiration for its Causal-Encoder Decoder (CED) transformer module.

Not Yet Confirmed

  • According to KOL MaxForAI's analysis of the V4.1 Flash paper's architecture diagrams (not officially confirmed), the model has at least three major architectural changes: a return to an encoder-decoder structure and the introduction of CSA2 cross-layer reuse, among others—what truly matters are the changes to the Transformer itself rather than parameters or benchmark scores.
  • Developer aHapBean argues that "Causal Encoder-Decoder" is a misnomer: V4.1 may still be a decoder-only LLM, and the name overstates the architectural innovation; a more accurate framing should follow the YoCo approach.

Why It Matters

  • @evijit points out that DeepSeek keeps innovating and open-sourcing at the architecture level rather than just scaling, raising the technical floor of the entire open-source community.
  • If practices like reusing KV caches across checkpoints hold up, they could significantly cut training and deployment costs, directly impacting inference economics.

Episode 10 · DeepSeek V4.1 Flash Architecture Breaks Down Asymmetric Design (2026-09-11, 2 posts)

DeepSeek's new technical report introduces DeepSeek-V4.1-Flash with an asymmetric causal encoder-decoder and CSA2 attention, cutting compute and memory costs while supporting a million-token context window.

Episode 11 · DeepSeek's New Model Report Highlights 4x Smaller KV Cache (2026-09-11, 3 posts)

Commentators reading DeepSeek's V4.1 Flash technical report highlight astonishing benchmark results, a 4x smaller KV cache than DSV4-Flash with major implications for inference memory and long-context costs, and more stable training.

Episode 12 · DeepSeek V4.1 Flash tops Vals open-source index at $0.30 per run (2026-09-11, 2 posts)

DeepSeek V4.1 Flash ranks first among open-weight models on the Vals Index, beating Kimi K3 while costing only $0.30 per run — roughly 1/20 the price of competitors like Kimi K3 and GLM 5.3.

Episode 13 · DeepSeek V4.1 Flash scores 40 in AA index despite rising hallucinations (2026-09-11, 2 posts)

Artificial Analysis gives DeepSeek V4.1 Flash a score of 40, beating DeepSeek V4 Pro though trailing GLM-5.3-Flash. Its AA-Omniscience rose 9 points to -5.3, but hallucination rates increased.

Episode 14 · Inside KV-Cache Sharing: How DeepSeek CED and GLM 5.2 Differ (2026-09-11, 4 posts)

Developers break down KV-cache sharing designs: YOCO lets all layers in the model's second half share one KV cache; DeepSeek's CED applies a similar YOCO-style global branch per CSA2 layer, caching only encoder KV; GLM 5.2 scores per token while sharing block indices for flexibility and performance.

Episode 15 · DeepSeek V4.1 Flash Deep Dive: KV Cache Compression Builds an Efficient Frontier Model (2026-09-11, 7 posts)

On 09-11, @nrehiew posted a long technical deep-dive thread on DeepSeek V4.1 Flash, with the throughline of how an "obsession with KV cache compression" produced an ultra-efficient flagship model. The series covers multimodal pretraining, optimizers and hyperparameters, training infrastructure, the attention indexer, global attention, and core architecture design.

Confirmed

  • Multimodal pretraining: pretrained on 45T image-text tokens, with a homegrown image encoder (a Siglip-style encoder for vision) and modality-level load balancing—which reminded the author of the original Chameleon approach
  • Optimizer: head-wise Muon including the vision model; sparse attention trained from scratch with no warmup; 196B-parameter Engram with Sinkhorn balancing
  • Training infrastructure: the key insight for the vision encoder is that features of one modality can be all-gathered in parallel while computing the other; context parallelism applied to ultra-long multi-image sequences
  • Attention indexer: the indexer itself is also sparse—the first full-mode indexer selects candidates, and later layers score them, forming a hierarchical filtering structure
  • Global attention: essentially upgraded compressed attention; all three variants first run the indexer on a small KV for selection, then concatenate with windowed KV before attention; they differ in how KV is reused—no reuse, reusing early-layer KV, etc.
  • Core architecture: a YOCO-style decoder-decoder design with 20 encoder layers + 20 decoder layers; addressing expensive prefill, the lower half generates the KV cache and shares it with the upper half via per-layer projections—effectively a way to halve the KV cache

Why it matters

The series systematically shows KV cache compression woven through every part of the model design—from architecture to attention mechanisms to indexer structure—a complete technical reference for understanding the engineering trade-offs behind this efficient flagship.