Leaked specs put Kimi K3 at 2.8T params — and why a 27B dense model isn't small

Xianbao_QIAN · x · 2026-08-16

XianbaoQIAN compiled an unverified comparison of rumored next-gen Chinese models (total / activated params):

In compute terms, the rumored Qwen 3.8 27B dense model is heavier than DeepSeek V4 Flash and comparable to MiniMax M3. The author stresses that 27B is not a small model, yet it now runs on consumer hardware — a testament to hardware and infra progress in recent years.

The quoted thread by @dashenwang explains why 27B is substantial: Google trained Gemma 3 27B on 14 trillion tokens using 6,144 TPU v5p chips. On the inference side alone, BF16 weights take 54GB, and 32K context with KV cache needs 72.7GB; training further requires gradients, optimizer states, activations and communication buffers. Reading 14T tokens at an aggressive 100k tokens per day, nonstop, would take one person 380,000 years.

So running 27B or even 70B locally is impressive not because PCs can train them, but because quantization and optimization squeeze a model trained on thousands of AI chips into a small box under your desk — "the process itself is pretty cyberpunk."

Original post →

More from Models

Models channel →