Model Parameters: Larger and Less Sparse
rasbt · x · 2026-07-16
The post shares a few observations regarding a specific model:
- The parameter count is 250B higher than GLM 5.2.
- It is less sparse than Kimi K2.5 1T: the latter has 3.2% sparsity and 32B active parameters, whereas this model has 4.2% sparsity and 41B active parameters.
- It does not use a hybrid architecture like Nemotron.
The author also expresses a strong desire to see a token/sec throughput comparison.
Related event: Kimi K3 Triggers a Reassessment of Chinese Frontier AI(94 posts)→
More from Models
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11