Predictions on the Next-Gen Sparse LLM Route
_xjdr · x · 2026-07-17
The author expresses no surprise at a certain TM approach, noting it aligns perfectly with what they've consistently recommended to Western labs.
Core predictions:
- Adopt an architecture like DSV3 shape, but with at least a 4:1 sliding window.
- Train with k2 rollouts, scaling parameters to at least 1T.
- Combine muon and mup, training on GB300.
- This is fundamentally an nmoe route.
They add that Meta should have taken this path from the start with Llama 4 and beyond. Looking ahead, they hope to see larger 3T+ models, higher expert sparsity, KV cache compression, base model releases, frontier RL refinement, and distillation papers. They believe that with higher sparsity, this architecture won't be the main bottleneck in the short term.
Related event: Kimi K3 Debuts Strong, Narrowing the Open-Weight Gap(184 posts)→
More from Models
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11