Is DeepSeek's rumored K3 a scaled-down model, or something bigger? X users debate
teortaxesTex · x · 2026-09-11
X users are speculating about DeepSeek's rumored next model, informally dubbed K3. One theory suggests it's a scaled-down model built on pre-V3.2 pretraining with dense MLA at every layer; teortaxesTex pushes back, arguing that if it were simply a smaller variant, DeepSeek would have just named it "K3-Flash" and everyone would be happy — implying the naming choice hints at something else. Unconfirmed rumor-level discussion; official details pending.
More from Models
- User Praises DeepSeek's Model as Surprisingly Fast and Good in Hands-on Test — MaziyarPanahi · 2026-09-11
- Sakana's Fugu Max routes a mixed model pool at $2/$6 per 1M tokens, claims tier-best benchmark score — SakanaAILabs · 2026-09-11
- Qwen3-8B gets a KV-approximation add-on that halves prefill time without touching the model — teortaxesTex · 2026-09-11
- Pro 20x tier burns 60% of weekly quota in under a day with GPT-6 Astra — rschu · 2026-09-11
- Google isn't honoring its own Gemini Grounded Search pricing: only 289 of 15,000+ requests counted as free — ItalyExpat · 2026-09-11
- DeepSeek update keeps cache hits mid-conversation, cuts costs 36.6% — teortaxesTex · 2026-09-11