Model Parameters: Larger and Less Sparse
rasbt · x · 2026-07-16
The post shares a few observations regarding a specific model:
- The parameter count is 250B higher than GLM 5.2.
- It is less sparse than Kimi K2.5 1T: the latter has 3.2% sparsity and 32B active parameters, whereas this model has 4.2% sparsity and 41B active parameters.
- It does not use a hybrid architecture like Nemotron.
The author also expresses a strong desire to see a token/sec throughput comparison.
Related event: Kimi K3 Triggers a Reassessment of Chinese Frontier AI(94 posts)→
More from Models
- A screenshot revisits GPT-4’s napkin-to-code demo and imagines GPT-N building GPT-N+1 — genmon · 2026-07-21
- Open weights, local models and open source are three different conversations — IronCuk · 2026-07-21
- Jack Clark says OpenAI’s internal-deployment safety notes help the whole frontier community — jackclarkSF · 2026-07-21
- Mindlab Research puts Macaron-V1-Venti on Hugging Face — External_Mood4719 · 2026-07-21
- ChatGPT often explains the wall before answering whether it is tilting — Aware-sky-3489 · 2026-07-21
- Grok website traffic rose 38.15% YoY to 736 million Q2 visits — XFreeze · 2026-07-21