DeepSeek appears to break its V<N> naming convention, new architecture said to train more stably
stochasticchasm · x · 2026-09-11
- A thread observes DeepSeek is breaking its own "V<N> = new pretrain" naming convention, despite having been one of the most consistent studios about version naming.
- The commenter reads the new release as a simplification of the V4 architecture, with the team claiming notably more stable training; they also find the inclusion of DeepSeek V1 in a KV cache comparison chart hilarious.
Related event: DeepSeek's New Model Report Highlights 4x Smaller KV Cache(3 posts)→
More from Models
- Astra pauses $200 Pro subscriptions as demand hits unprecedented levels — emax · 2026-09-11
- inclusionAI's Open-Source Ling-3.0-flash-VL Multimodal Model Trends on Hugging Face — inclusionAI · 2026-09-11
- Tiny KV Cache via Shared Global KV Plus Per-Layer SWA? New Architecture Speculation — stochasticchasm · 2026-09-11
- DeepSeek v4.1 Flash Tested Across 8 Coding Harnesses: Performs Best in Minimal Setups — mariofilhoml · 2026-09-11
- Would ChatGPT Plus users accept a 24-hour usage limit instead of weekly caps? — SuaveSteve · 2026-09-11
- Leaked Qwen next-gen model shows record n-gram params, first two layers SWA-only — stochasticchasm · 2026-09-11