Report: DeepSeek views Kimi's approach as hard to scale, seeks a cheap scalable recipe

teortaxesTex · x · 2026-10-03

Responding to whether DeepSeek and Zhipu hit pretraining walls, teortaxesTex claims DeepSeek considers Kimi's approach worse: the 2.8T model can't be served at sufficient volume at this stage given its compute-heavy architecture, and Zhipu genuinely struggles with scaling. He says DeepSeek wants a recipe that scales while staying cheap. Unverified secondhand claim.

Original post →

More from Models

Models channel →