Discussion on LatentMoE Sparsity and Scaling Efficiency
ivan_bezdomny · x · 2026-07-17
The thread focuses on sparsity designs like LatentMoE:
- The author has long argued that activation ratios can drop below 2%, with reasons to sustain scaling even at <1%.
- The cited text mentions that LatentMoE activates only 16/896 experts. Combined with Kimi Delta Attention and AttnRes, it reportedly achieves 2.5x more efficient scaling.
- It also notes that rumors of OpenAI and Anthropic moving towards significantly higher sparsity align with this trajectory.
- The poster adds a cost comparison: K3 has a per-token cost similar to V4, but being much larger, its overall cost is 17x higher, raising questions about profit margins.
More from Models
- Gemini 3.6 Flash appears live in Studio with $1.50 input pricing — ivan_bezdomny · 2026-07-21
- Artificial Analysis ranks Gemini 3.6 Flash at 50 on its updated intelligence index — Angaisb_ · 2026-07-21
- Google appears to have quietly shipped Gemini 3.6 Flash, with lower pricing and better agentic scores — xiaohu · 2026-07-21
- Google ships three more Gemini variants while 3.5 Pro slips again — Miserable-Archer-631 · 2026-07-21
- Google Quietly Launches Gemini 3.6 Flash: Cheaper, Stronger, and Agentic-Focused — OwariDa · 2026-07-21
- A user says 10–12 hours with Claude equals 3–4 hours with Grok Build — Daniel_Farinax · 2026-07-21