FULL STORY
DeepSeek V4.1 Flash: From Limited Beta to Open-Source Release
After a limited beta on Sept 8, DeepSeek officially released and open-sourced V4.1 Flash on Sept 10 with price cuts. Benchmarks show near-frontier performance, a 4x smaller KV cache, and the top spot on the Vals open-weight leaderboard.
2026-09-08 ~ 2026-09-11 · 15 episodes · 154 posts
Episode 1 · DeepSeek V4.1 Flash Opens Limited Internal Beta with New Architecture and Native Multimodality (2026-09-08, 18 posts)
Starting September 8, multiple sources reported that DeepSeek quietly opened a limited internal beta for V4.1 Flash, a "middle version" with model ID deepseek-v4.1-flash-expires-on-0910, the slug indicating expiry on September 10. Key upgrades include a brand-new architecture and native multimodality, with the official line claiming stronger capability, faster speed, and lower cost. Calling conventions, billing, and concurrency rules are now largely clear, developer benchmarks have emerged, and the model has appeared on Hugging Face, though an official public announcement and full evaluations remain pending.
Confirmed
- Jiqizhixin reported that DeepSeek officially announced the internal test of the V4.1 Flash middle version; APPSO, Zhidongxi and other outlets confirmed the same.
- Usage: keep baseurl unchanged and set the model name to an ID beginning with deepseek-v4.1-flash-expires-on-09 (e.g., deepseek-v4.1-flash-expires-on-0910), where the date marks expiry.
- According to vista8's relay of an official WeChat group notice, billing matches V4 Flash with max 20 concurrent requests per account; teortaxesTex's leaked info points to the same concurrency limit and a vision-capable build (V4-Flash-Vision [Intermediate]).
- On pricing, MiaAIlab (relayed by MicahBerkley) described the price as extremely aggressive: roughly 58 million tokens per US dollar, available for about 2 days.
- Third-party account LuminaBench spotted the intermediate build being tested via API; tokenbender suggested it precedes the official release.
- On September 10, Kevin Kern tested the preview endpoint, measuring 300 tok/s and heavily parallelizing subagents; he noted the preview endpoint does not yet accept vision input. Zhidongxi reported output speeds up to 507 tokens/s.
- On September 10, Reddit users found a deepseek-ai/DeepSeek-V4.1-Flash model page on Hugging Face, suggesting a quiet official upload; specs unconfirmed.
Unconfirmed
- The news initially spread via bloggers such as op7418 and AGI Hunt from a DeepSeek WeChat group, unverified at the time; Jiqizhixin later framed it as officially announced, but a formal public announcement and full capability details are still pending.
- @NFTChen (relayed by solyarisoftware) claimed V4.1 Flash would be officially released around September 10 Beijing time, with V4 Pro requests routed at Flash pricing; not confirmed by DeepSeek.
- The preview endpoint's lack of vision input sits in tension with the "native multimodality" claim; whether the formal release adds vision remains to be seen.
- teortaxesTex relayed community banter hoping it can at least beat GLM 5.3 Flash outside of vision—opinion only, no benchmark basis.
Why it matters
- This is DeepSeek's first reported use of an "expires-on" middle-version limited beta with explicit billing and concurrency rules plus aggressive pricing, signaling a product cadence of rapid iteration through real developer workloads before formal release.
- Native multimodality marks a major upgrade to DeepSeek's flagship line; with 300 tok/s measured, an appearance on Hugging Face, and the community benchmarking it against GLM 5.3 Flash, the formal release's multimodal capabilities deserve close attention.
- How to call DeepSeek V4.1 Flash beta: model name and rate limits revealed — 赛博禅心 · 2026-09-08
- Rumored DeepSeek V4.1 Flash surfaces, user hopes it beats GLM 5.3 Flash — teortaxesTex · 2026-09-08
- DeepSeek V4.1 Flash mid-cycle build enters beta with native multimodal support — AGI Hunt · 2026-09-08
- Rumor: DeepSeek V4.1 Flash released with stronger native multimodal — op7418 · 2026-09-08
- DeepSeek V4.1 Flash reportedly in private beta with new architecture and native multimodal support — vista8 · 2026-09-08
- DeepSeek opens internal beta of V4.1 Flash with native multimodal support — jiqizhixin · 2026-09-08
- DeepSeek spotted testing V4-Flash-Vision: new arch, native multimodal, same price — teortaxesTex · 2026-09-08
- DeepSeek V4.1 Flash beta goes live with new architecture, native multimodality, up to 507 tokens/s — 智东西 · 2026-09-08
- DeepSeek quietly tests V4.1 Flash API beta with new architecture and native multimodality — tokenbender · 2026-09-08
- DeepSeek Opens Limited-Time Beta of V4.1Flash: New Architecture, Native Multimodal, Expires Sept 10 — APPSO · 2026-09-08
- DeepSeek V4.1 Flash interim build reportedly in beta: native multimodal, faster and cheaper — aigclink · 2026-09-08
- DeepSeek Flash 4.1 spotted testing via API with new architecture, release imminent — kimmonismus · 2026-09-08
- DeepSeek V4.1 Flash in internal beta: native multimodal, same price as V4 Flash — Nunki08 · 2026-09-08
- DeepSeek v4.1 Flash flash sale: 58M tokens for $1, available for 2 days only — MicahBerkley · 2026-09-09
- DeepSeek v4.1 flash preview spotted running at ~300 tok/s, vision still missing — kevinkern · 2026-09-10
- DeepSeek v4.1 Flash preview spotted: ~300 tok/s, fires lots of subagents, no vision yet — kevinkern · 2026-09-10
- Rumor: DeepSeek V4.1 Flash to launch ~Sept 10, V4 Pro requests rerouted at Flash pricing — solyarisoftware · 2026-09-10
- DeepSeek-V4.1-Flash spotted on Hugging Face, unannounced — t4a8945 · 2026-09-10
Episode 2 · DeepSeek V4.1 Flash Tested: Blazing 350 Tokens/s but Still Experimental (2026-09-08, 2 posts)
Hands-on tests of DeepSeek's experimental V4.1-Flash model show decoding speeds of roughly 300-400 tokens per second at competitive prices, though reviewers note it remains early-stage and prone to overthinking.
- DeepSeek V4.1-Flash hands-on: 350 t/s decoding speed but still very experimental — teortaxesTex · 2026-09-08
- DeepSeek V4.1 Flash tested: 300-400+ tok/s, cheap and fast, but prone to overthinking — WorldofAI · 2026-09-09
Episode 3 · DeepSeek cuts V4-Flash API prices with new peak/off-peak billing from Sept 10 (2026-09-08, 6 posts)
DeepSeek officially announced a price cut for its V4-Flash (flash) API series effective Sept 10, 12:00 Beijing time, restoring input prices to pre-hike levels and switching to peak/off-peak billing, with the largest discounts in off-peak hours. The new scheme replaces the expiring "Intermediate" V4.1 pricing; the price sheet had leaked via community channels a day earlier.
Confirmed
- DeepSeek platform announcement: V4-Flash price cut from Sept 10, 12:00 (reported by 机器之心, 量子位, aigclink).
- Off-peak pricing (yuan per million tokens): cache-hit input 0.02 (down 60% from 0.05); cache-miss input 1 (down 33.33% from 1.5); output 4 (down 11.11% from 4.5).
- Peak hours are weekdays 9:00–12:00 and 14:00–18:00; the rest is off-peak. Peak prices are higher, per the official schedule.
- The new scheme replaces the expiring "Intermediate" V4.1 pricing (per teortaxesTex).
- michaelsoftbinbows noted that DeepSeek V4 Flash (0731) via openinference on OpenRouter costs only $0.05 input / $0.16 output per 1M tokens off-peak, below official pricing (from $0.22), drawing wide discussion.
Unconfirmed
- WeChat account 赛博禅心 mentioned a 20-concurrency limit accompanying the new scheme, flagged as pending official confirmation; the official announcement is silent on this.
Timeline
- Sept 8: teortaxesTex summarized the leaked pricing; 赛博禅心 posted the cut with effect time and concurrency details, unconfirmed.
- Sept 9: Official announcement confirmed via 量子位, 机器之心, aigclink; OpenRouter third-party low pricing sparked discussion.
- Sept 10, 12:00: New prices take effect.
Why it matters
- Input prices returning to pre-hike levels partially rolls back the earlier increase, directly benefiting developers and vendors relying on DeepSeek APIs.
- Off-peak cache-hit pricing as low as 0.02 yuan per million tokens, combined with time-based billing, will push users toward off-peak calls and higher cache utilization, reflecting shifts in DeepSeek's capacity and scheduling.
- The notable gap between official and third-party channels (e.g., OpenRouter) offers developers a cheaper alternative access path.
- DeepSeek cuts flash-series prices to 1 yuan/M input tokens, launches v4.1-flash — 赛博禅心 · 2026-09-08
- DeepSeek's Post-Sept 10 Pricing: Premium Output, Off-Peak Cache Hits at ¥0.02 — teortaxesTex · 2026-09-08
- DeepSeek cuts flash pricing up to 60% as v4.1 flash quietly goes live — 量子位 · 2026-09-09
- DeepSeek cuts flash pricing up to 60%, introduces peak/off-peak billing — aigclink · 2026-09-09
- DeepSeek V4 Flash listed at $0.05/$0.16 per 1M off-peak on OpenRouter, 4x below official pricing — michaelsoft__binbows · 2026-09-09
- DeepSeek cuts V4-Flash API prices Sept 10, inputs back to pre-hike levels as V4.1 test model appears — 机器之心 · 2026-09-09
Episode 4 · DeepSeek V4.1 Flash tested across tasks: near-frontier performance at a fraction of the cost (2026-09-09, 15 posts)
DeepSeek V4.1 Flash, along with the V4.1 Flash Vision Beta released the same day, has been put through multiple rounds of independent testing in the days since launch. The consensus: its performance approaches or even matches frontier closed-source models in some areas, at just 1/100 to 1/30 of the cost.
Confirmed
- Design benchmark: In a comparison of everyday design tasks, blogger NFTChen scored DeepSeek V4.1 Flash at 81.2—98% of GPT-6 Astra's 82.7 and ahead of Claude Fable 5.1's 80.3—at roughly 1.4% of Astra's cost; OpenDesign Arena leaderboard data shows similar value.
- Cybersecurity benchmark: The pilvar222 team found V4.1 Flash rediscovers 65.6% of recent CVEs in the benchmark in a single run (vs. 55.2% for the old version), with pass@3 catching 84.4% (vs. 75%), matching frontier models.
- Real-world bug fixing: Pawel Huryn tested with 2 repos and 105 hidden real bugs; V4.1 Flash landed on the Pareto frontier—for reference, Opus 5 (max) found 27 bugs at $51.33, while V4.1 Flash costs about 3.5% of Opus (1/28 the price). The author also added an FAQ addressing common objections.
- Vision: On launch day, Reddit user cheezeerd compared V4.1 Flash Vision Beta with V4 Flash Vision across 5 vision tasks, with full video walkthroughs of the first draft and three revisions—concluding the new version produces noticeably better, more reliable output at lower API pricing.
Why it matters
- Tests from multiple independent sources across different task types (design, security, code, vision) corroborate V4.1 Flash's value proposition, rather than resting on a single benchmark.
- If the results replicate, ultra-cheap near-frontier models will further squeeze pricing for closed-source flagships and lower the barrier for developers and enterprises to switch.
Not yet confirmed
- All results come from third-party individuals or teams with varying samples and methodologies; official metrics and large-scale replications have yet to appear.
- DeepSeek V4.1 Flash Vision Beta tested in 5 visual tasks: far better and cheaper than V4 — cheezeerd · 2026-09-09
- DeepSeek V4.1 Flash Vision Beta tested: big quality jump and lower API pricing — cheezeerd · 2026-09-09
- DeepSeek V4.1 Finds 65.6% of CVEs on Cybersecurity Bench, Matches Frontier Models at a Fraction of the Cost — teortaxesTex · 2026-09-09
- DeepSeek v4.1 Flash Hits 98% of Astra's Score at 1.4% of the Cost — uxl · 2026-09-09
- Blogger's test: DeepSeek V4.1 Flash scores 98% of GPT-6 Astra at 1.4% of the cost — solyarisoftware · 2026-09-10
- 105 hidden bugs tested: DeepSeek V4.1 Flash near Opus 5 at 1/28th the cost — PawelHuryn · 2026-09-10
- DeepSeek V4.1 Flash fixes 24/105 real bugs for $1.80 vs Opus 5's $51.33 — PawelHuryn · 2026-09-10
- Hidden-bug hunt: DeepSeek V4.1 Flash fixes 24/105 bugs at $1.80 vs Opus 5's $51.33 — soveLight · 2026-09-10
- DeepSeek V4.1-Flash beats V4-Pro on Terminal-Bench at 1/20th the cost of Opus 5 — dejavucoder · 2026-09-10
- 105 Hidden Bugs Benchmark: DeepSeek V4.1 Flash Scores 24 at $1.80 vs Opus 5's $51 — sanderjson · 2026-09-10
- DeepSeek-V4.1-Flash aces 22 knowledge-work tasks for just $0.34 in eval run — realsohamparekh · 2026-09-10
- Third-Party Benchmark Claims DeepSeek V4.1 Flash Matches 98% of GPT-6 at 1.4% of Cost — ChrisUniverse · 2026-09-10
- Internal eval: DeepSeek V4.1 Flash nearly SOTA at 98% lower cost than Fable 5.1 — himanshustwts · 2026-09-11
- DeepSeek 4.1 Flash Hits 200 TPS Full Precision on 4 Max-Qs with Just 64GB RAM — TheZachMueller · 2026-09-11
- DeepSeek V4.1 Flash scores 40 on Intelligence Index at 190 tokens/sec — ArtificialAnlys · 2026-09-11
Episode 5 · DeepSeek releases V4.1-Flash: 552B MoE beats flagships at low cost (2026-09-09, 56 posts)
On September 10, DeepSeek officially released and open-sourced its new model V4.1 Flash, rolling out simultaneously on web, app, and API, with the API model named deepseek-flash. This is one of the most notable new open-source model releases right now: performance reportedly surpasses their own flagship V4 Pro across the board, at a sharply lower price.
Confirmed
- Architecture and scale: a 552B total-parameter asymmetric MoE model with a brand-new Causal-Encoder-Decoder architecture, activating only 8B parameters on the input side and 16B on the output side; supports up to 1 million tokens of context, with KV cache overhead cut by 7/8.
- Benchmarks: per @ChrisGPT, V4.1 Flash beats GPT-5.6 Sol and Opus 5 on multiple agentic/coding benchmarks, e.g., CyberGym 88.1 vs 84.5; input pricing as low as $0.14/M.
- V4 Pro retirement and routing: the official announcement states that since V4.1 Flash surpasses V4 Pro across performance, cost, speed, and total time, all V4 Pro requests will be automatically routed to V4.1 Flash and billed at the lower price until V4.1 Pro launches. Several users (e.g., @gaganghotra) have already received deprecation notices.
- Older models sunset: the official API announcement discloses that legacy models like V4-Flash and V4-Flash-Vision-E will be retired.
Why it matters
- V4.1 Flash matches and exceeds flagship V4 Pro performance with far fewer activated parameters and much lower cost—combined with its open-source release, this could further drive down industry inference prices and pressure closed-model vendors.
- V4 Pro users should watch for behavior changes from the automatic routing; developers relying on its capabilities should evaluate V4.1 Flash's actual performance before switching.
- The 1M-token context and 7/8 KV cache savings have direct implications for usability and cost in long-document and agentic workflow scenarios.
- DeepSeek routes V4 Pro requests to faster, cheaper V4.1 Flash, hints V4.1 Pro is coming — zephyr_z9 · 2026-09-09
- DeepSeek releases V4.1 Flash: 22% cheaper than V4 and better performance — deliprao · 2026-09-09
- DeepSeek ships V4.1 Flash: 552B MoE on new architecture, MIT-licensed — deepseek-ai · 2026-09-10
- DeepSeek releases V4.1 Flash: 552B MoE beats V4 Pro, cuts prices, retires flagship — DeepSeek · 2026-09-10
- DeepSeek Launches V4.1-Flash, Smallest Model in New Architecture Family with Native Vision — deepseek_ai · 2026-09-10
- DeepSeek Details V4.1-Flash: 552B MoE with Asymmetric 8B/16B Active Params, Beats Its Flagship — deepseek_ai · 2026-09-10
- DeepSeek-V4.1-Flash Hits the API: 1/8 the SSD Cache, V4-Pro Being Phased Out — deepseek_ai · 2026-09-10
- DeepSeek Open-Sources V4.1-Flash Under MIT License, Cuts Off-Peak API Rates to 50% — deepseek_ai · 2026-09-10
- DeepSeek V4.1 Flash drops: 552B MoE activating just 8B, beating GPT-5.6 Sol at $0.14/M — ChrisGPT · 2026-09-10
- DeepSeek to retire v4 Pro, reroute all requests to new v4.1 Flash — gaganghotra_ · 2026-09-10
- DeepSeek releases V4.1 Flash: 552B MoE, 1M context, SGLang day-0 support — BanghuaZ · 2026-09-10
- Inside V4.1 Flash: Per-Token KV Cache Down to 890 Bytes, Peak-Valley Pricing — 赛博禅心 · 2026-09-10
- DeepSeek Open-Sources V4.1 Flash: 552B Asymmetric MoE Replaces V4 Pro — 机器之心 · 2026-09-10
- DeepSeek-V4.1-Flash appears on Hugging Face: MIT-licensed multimodal model with FP8 weights — AIFlow_ML · 2026-09-10
- DeepSeek V4.1 Flash weights go live on Hugging Face under MIT license — solyarisoftware · 2026-09-10
- DeepSeek releases V4.1-Flash, smallest of new architecture family with native vision — eliebakouch · 2026-09-10
- DeepSeek launches V4.1-Flash: smallest model in new architecture family with native vision — reach_vb · 2026-09-10
- DeepSeek Launches V4.1-Flash: Smallest Model in New Architecture Family with Native Vision — teortaxesTex · 2026-09-10
- DeepSeek launches V4.1-Flash, smallest model in new architecture family with native vision — basedjensen · 2026-09-10
- DeepSeek V4.1 Flash now live on App, Web and API with open weights — gaganghotra_ · 2026-09-10
Episode 6 · DeepSeek V4.1 Flash Leak: 552B Asymmetric MoE Reportedly Rivals GPT-5.6 (2026-09-10, 20 posts)
On September 10, several bloggers collectively leaked what are claimed to be official benchmark results and technical details for DeepSeek V4.1 and V4.1 Flash; none of the information has been officially confirmed.
Confirmed
- All information is unverified rumor, mainly from bloggers such as teortaxesTex, eliebakouch, and ChrisGPT, whose posts corroborate each other.
- The model has 552B total parameters with a brand-new Causal-Encoder-Decoder architecture, 8B input activation and 16B output activation.
- Leaked scores: TerminalBench 4.0 at 31.2, CyberGym 88.1 (vs. 84.5 for GPT-5.6 Sol), HLE 63.9; ChrisGPT claims it beats GPT-5.6 Sol on multiple agentic/coding benchmarks.
- The technical report cited by eliebakouch shows V4.1 Flash is a natively vision-capable model trained on 45T tokens, with an architecture praised as the most novel in years.
Architecture and Engineering Highlights
- KV cache compressed to 890 bytes/token, under 1GB for a million tokens; compared with V4-Flash, HBM requirements are down another 3.9x and SSD requirements down 8x, still unsurpassed half a year after release (per teortaxesTex).
- According to teortaxesTex, V4.1 removes V4's stopgap patches: dropping HCA, generalizing CSA into a more universal primitive, eliminating the warmup stage of sparse attention, and adopting a simpler single-pass mHC.
Not Yet Confirmed
- All scores, specs, and architecture details lack official sources; the final release may differ.
Why It Matters
- If true, V4.1 directly challenges GPT-5.6 Sol on agentic/coding capability while slashing deployment memory requirements, with significant implications for inference costs in an open-source ecosystem.
- Leaked DeepSeek V4.1 benchmarks show 552B new-architecture model hitting 63.9 HLE with tools — teortaxesTex · 2026-09-10
- DeepSeek V4.1 Leak: 552B New Architecture, HBM Needs Cut 3.9x, Huge Benchmark Gains — teortaxesTex · 2026-09-10
- DeepSeek slashes KV cache to 890 bytes/token, hinting at million-agent swarms — teortaxesTex · 2026-09-10
- DeepSeek V4.1 Flash rumored: 552B MoE with asymmetric CED architecture, beats flagship — _AndrewZhao · 2026-09-10
- DeepSeek V4.1 Flash benchmarks leak: beats GPT-5.6 Sol on agentic/coding tasks at $0.14/M input — ChrisGPT · 2026-09-10
- DeepSeek V4.1 Flash leak: 552B MoE with 8/16B active, native vision, praised as most novel arch in years — eliebakouch · 2026-09-10
- Leak: DeepSeek V4.1 reportedly drops V4's HCA, generalizes CSA, removes sparse-attention warmup — zephyr_z9 · 2026-09-10
- DeepSeek V4.1 Flash hits 74.2 on DeepSWE, beating Opus 5 and Gemini 3.8 Flash; asymmetric MoE revealed — nrehiew_ · 2026-09-10
- DeepSeek-V4.1-Flash reportedly offers continuous reasoning effort from 1 to 100 — zainhas · 2026-09-10
- DeepSeek V4.1-Flash leak: 552B params claiming GPT-5.6-level scores via CED architecture — iScienceLuvr · 2026-09-10
- DeepSeek-V4.1-Flash paper hints at continuously tunable reasoning effort from 1 to 100 — zainhas · 2026-09-10
- Rumored DeepSeek V4.1 Flash details point to asymmetric-activation MoE with much lower cost — gaganghotra_ · 2026-09-10
- DeepSeek-V4.1-Flash offers continuously controllable reasoning effort from 1 to 100 — zainhas · 2026-09-10
- 552B params but only 8B active: DeepSeek V4.1 Flash's wild efficiency numbers — Scobleizer · 2026-09-10
- DeepSeek V4.1-Flash: KV cache 437x smaller, ~40x cheaper than Claude Opus 4.8 — AdinaYakup · 2026-09-10
- Leak claims DeepSeek v4.1 Flash matches or beats GPT-5.6 Sol and Claude Opus 5 — airesearch12 · 2026-09-10
- DeepSeek-V4.1-Flash leaks: 552B backbone activating just 8B, 8x less KV cache — bodonoghue85 · 2026-09-10
- DeepSeek V4.1 Flash rumored: 552B params, new arch with YOCo KV compression — donglixp · 2026-09-10
- DeepSeek V4.1-Flash turns heads: GPT-5.6-level benchmarks at 552B params — teortaxesTex · 2026-09-10
- DeepSeek V4.1 Flash weights out, outperforming Opus-5 and GPT-5.6-Sol in key areas — Sentdex · 2026-09-11
Episode 7 · DeepSeek V4.1 Flash Benchmark Results Leak, Sparking Community Discussion (2026-09-10, 2 posts)
Early benchmark screenshots circulating on Reddit show DeepSeek V4.1 Flash achieving SOTA on the deepswe leaderboard alongside sol and fable5, with competitive terminal bench results, sparking community discussion.
- DeepSeek v4.1 flash evals leak: SOTA on deepswe, competitive on terminal bench — zainhas · 2026-09-10
- DeepSeek V4.1 Flash benchmark chart circulates after release — toastisthicc · 2026-09-10
Episode 8 · Bug Hunt Bench Tests 105 Real Bugs; DeepSeek V4.1 Flash Stands Out on Cost (2026-09-10, 8 posts)
Pawel Huryn (author of The Product Compass) launched and maintains the Bug Hunt Bench benchmark: 105 real bugs were deliberately planted in two production codebases, and frontier models including GPT-6, Claude, Grok, Gemini, DeepSeek, Kimi, and GLM were tasked with fixing them in their respective agentic configurations, ranked by blind review of the number of fixes, with only one prompt given per repository, a maximum score of 105, and results continuously updated on its GitHub leaderboard page.
Confirmed
- GPT-6 Astra supplementary results: low tier 27/105, medium tier 34/105; these were previously released on September 4-5 and can be compared via filters on the leaderboard page.
- DeepSeek-V4.1-Flash (high tier): fixed 24/105 at max effort, with a total cost of only $1.80 and a runtime of 42.6 minutes.
Why it matters
- With its setup of pre-planted bugs in real repositories, blind review, and a single prompt, this benchmark is closer to real-world bug-fixing scenarios than traditional synthetic benchmarks.
- Huryn highlighted that DeepSeek-V4.1-Flash fixed 24 bugs for roughly $1.80, delivering standout cost-effectiveness in its price class and providing valuable selection data for budget-conscious teams.
- Bug Hunt Bench grades frontier models on 105 real bugs; DeepSeek-V4.1-Flash lands 24/105 for $1.80 — PawelHuryn · 2026-09-10
- Bug Hunt Bench adds GPT-6 Astra scores: 27/105 low, 34/105 medium — PawelHuryn · 2026-09-10
- Same Bug Benchmark: GPT-6 Astra Medium Fixes 34/105, Low Scores 27 — PawelHuryn · 2026-09-10
- DeepSeek V4.1 Flash fixes 24 of 105 hidden bugs at $1.80, near Opus-level at 1/28th the cost — AlexSed90 · 2026-09-10
- DeepSeek V4.1 Flash scores 24/105 for $1.80 in community eval test — PawelHuryn · 2026-09-10
- Bug Hunt Bench ranks frontier coding models on 105 real planted bugs, with 200x cost spread — PawelHuryn · 2026-09-10
- DeepSeek V4.1 Flash beats V4 Pro in retest: 24/105 vs 16/105 at lower cost — PawelHuryn · 2026-09-10
- Bug Hunt Bench: Frontier Coding Models Graded Blind on 105 Planted Real-Repo Bugs — PawelHuryn · 2026-09-10
Episode 9 · Inside DeepSeek V4.1 Flash: YOCO at its core and KV cache reuse (2026-09-10, 7 posts)
On September 10, several bloggers focused on architectural changes in DeepSeek's new V4.1 model: DeepSeek is said to be returning to an encoder-decoder architecture, with multiple innovations at the architectural level and full open-sourcing—widely seen as continually raising the technical floor of the AI research ecosystem. Most of this information comes from paper details and KOL analysis; the official position is not yet unified, and the architecture's naming remains disputed.
Confirmed
- DeepSeek has open-sourced the technologies related to the new architecture.
- Astra Pro's deep dive into the V4.1 paper found that DeepSeek reuses KV caches across different checkpoints, prompting commenters to call it a "strange model factory."
- DeepSeek cites the 2024 YoCo (You Only Cache Once) paper as the main inspiration for its Causal-Encoder Decoder (CED) transformer module.
Not Yet Confirmed
- According to KOL MaxForAI's analysis of the V4.1 Flash paper's architecture diagrams (not officially confirmed), the model has at least three major architectural changes: a return to an encoder-decoder structure and the introduction of CSA2 cross-layer reuse, among others—what truly matters are the changes to the Transformer itself rather than parameters or benchmark scores.
- Developer aHapBean argues that "Causal Encoder-Decoder" is a misnomer: V4.1 may still be a decoder-only LLM, and the name overstates the architectural innovation; a more accurate framing should follow the YoCo approach.
Why It Matters
- @evijit points out that DeepSeek keeps innovating and open-sourcing at the architecture level rather than just scaling, raising the technical floor of the entire open-source community.
- If practices like reusing KV caches across checkpoints hold up, they could significantly cut training and deployment costs, directly impacting inference economics.
- DeepSeek returns to encoder-decoder with new open architecture changes — evijit · 2026-09-10
- DeepSeek's V4.1 Flash reworks the Transformer again: asymmetric encoder-decoder and CSA2 — evijit · 2026-09-10
- DeepSeek cites 2024 YoCo paper as inspiration behind its CED transformer blocks — jm_alexia · 2026-09-10
- DeepSeek's V4.1 paper reveals KV cache reuse across checkpoints — 'a very strange model factory' — teortaxesTex · 2026-09-10
- DeepSeek V4.1 'causal encoder-decoder' label debated: likely a YOCO-style KV-reuse design — donglixp · 2026-09-10
- DeepSeek V4.1 Flash reportedly adopts YOCO architecture for order-of-magnitude KV savings — donglixp · 2026-09-11
- DeepSeek researcher: V4.1 Flash builds on YOCO as architecture innovation enters 'second half' — donglixp · 2026-09-11
Episode 10 · DeepSeek V4.1 Flash Architecture Breaks Down Asymmetric Design (2026-09-11, 2 posts)
DeepSeek's new technical report introduces DeepSeek-V4.1-Flash with an asymmetric causal encoder-decoder and CSA2 attention, cutting compute and memory costs while supporting a million-token context window.
- DeepSeek-V4.1-Flash architecture dissected: asymmetric causal encoder-decoder, CSA2 attention, Engram memory — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Architecture: 552B MoE with Asymmetric 8B Read / 16B Decode Compute — demian_ai · 2026-09-11
Episode 11 · DeepSeek's New Model Report Highlights 4x Smaller KV Cache (2026-09-11, 3 posts)
Commentators reading DeepSeek's V4.1 Flash technical report highlight astonishing benchmark results, a 4x smaller KV cache than DSV4-Flash with major implications for inference memory and long-context costs, and more stable training.
- Blogger flags new model's standout tech report: high benchmarks and 4x smaller KV cache vs dsv4-flash — stochasticchasm · 2026-09-11
- DeepSeek's New Model: 4x Smaller KV Cache Than DSV4-Flash and More Stable Training — stochasticchasm · 2026-09-11
- DeepSeek appears to break its V<N> naming convention, new architecture said to train more stably — stochasticchasm · 2026-09-11
Episode 12 · DeepSeek V4.1 Flash tops Vals open-source index at $0.30 per run (2026-09-11, 2 posts)
DeepSeek V4.1 Flash ranks first among open-weight models on the Vals Index, beating Kimi K3 while costing only $0.30 per run — roughly 1/20 the price of competitors like Kimi K3 and GLM 5.3.
- DeepSeek V4.1 Flash tops Vals open-weight index at $0.30 per test, with the smallest skills gap — teortaxesTex · 2026-09-11
- DeepSeek v4.1 flash reportedly tops Vals Intelligence Index at 1/20th the cost of rivals — zainhas · 2026-09-11
Episode 13 · DeepSeek V4.1 Flash scores 40 in AA index despite rising hallucinations (2026-09-11, 2 posts)
Artificial Analysis gives DeepSeek V4.1 Flash a score of 40, beating DeepSeek V4 Pro though trailing GLM-5.3-Flash. Its AA-Omniscience rose 9 points to -5.3, but hallucination rates increased.
- DeepSeek-V4.1-Flash scores 40 on Artificial Analysis Index, beating DeepSeek-V4-Pro — scaling01 · 2026-09-11
- DeepSeek V4.1 Flash evals: accuracy up but non-hallucination rate falls — ArtificialAnlys · 2026-09-11
Episode 14 · Inside KV-Cache Sharing: How DeepSeek CED and GLM 5.2 Differ (2026-09-11, 4 posts)
Developers break down KV-cache sharing designs: YOCO lets all layers in the model's second half share one KV cache; DeepSeek's CED applies a similar YOCO-style global branch per CSA2 layer, caching only encoder KV; GLM 5.2 scores per token while sharing block indices for flexibility and performance.
- YOCO explained: one shared KV cache reused across the model's second half — stochasticchasm · 2026-09-11
- DeepSeek's CED applies YOCO-style shared KV caching only to global CSA2 branches — stochasticchasm · 2026-09-11
- DeepSeek's CED vs GLM's KV reuse: a developer unpacks how the cache-sharing modes actually differ — stochasticchasm · 2026-09-11
- GLM 5.2's KV cache sharing modes explained: blockwise to tokenwise scoring — stochasticchasm · 2026-09-11
Episode 15 · DeepSeek V4.1 Flash Deep Dive: KV Cache Compression Builds an Efficient Frontier Model (2026-09-11, 7 posts)
On 09-11, @nrehiew posted a long technical deep-dive thread on DeepSeek V4.1 Flash, with the throughline of how an "obsession with KV cache compression" produced an ultra-efficient flagship model. The series covers multimodal pretraining, optimizers and hyperparameters, training infrastructure, the attention indexer, global attention, and core architecture design.
Confirmed
- Multimodal pretraining: pretrained on 45T image-text tokens, with a homegrown image encoder (a Siglip-style encoder for vision) and modality-level load balancing—which reminded the author of the original Chameleon approach
- Optimizer: head-wise Muon including the vision model; sparse attention trained from scratch with no warmup; 196B-parameter Engram with Sinkhorn balancing
- Training infrastructure: the key insight for the vision encoder is that features of one modality can be all-gathered in parallel while computing the other; context parallelism applied to ultra-long multi-image sequences
- Attention indexer: the indexer itself is also sparse—the first full-mode indexer selects candidates, and later layers score them, forming a hierarchical filtering structure
- Global attention: essentially upgraded compressed attention; all three variants first run the indexer on a small KV for selection, then concatenate with windowed KV before attention; they differ in how KV is reused—no reuse, reusing early-layer KV, etc.
- Core architecture: a YOCO-style decoder-decoder design with 20 encoder layers + 20 decoder layers; addressing expensive prefill, the lower half generates the KV cache and shares it with the upper half via per-layer projections—effectively a way to halve the KV cache
Why it matters
The series systematically shows KV cache compression woven through every part of the model design—from architecture to attention mechanisms to indexer structure—a complete technical reference for understanding the engineering trade-offs behind this efficient flagship.
- DeepSeek V4.1 Flash notes: how obsessing over KV cache compression yields a hyper-efficient frontier model — nrehiew_ · 2026-09-11
- DeepSeek V4.1 Flash multimodal: 45T image-text tokens and modality-level load balancing — nrehiew_ · 2026-09-11
- DeepSeek V4.1 Flash halves KV cache with YOCO-style decoder-decoder design — nrehiew_ · 2026-09-11
- DeepSeek V4.1 Flash's global attention: compressed attention with three KV-reuse variants — nrehiew_ · 2026-09-11
- DeepSeek V4.1 Flash: even the sparse indexer is itself sparse, with hierarchical candidate selection — nrehiew_ · 2026-09-11
- DeepSeek V4.1 Flash uses Muon optimizer and Sinkhorn balancing for 196B-param Engram — nrehiew_ · 2026-09-11
- DeepSeek V4.1 Flash infra: Siglip-style vision encoder and shadow indexer workers — nrehiew_ · 2026-09-11