Qwen teases new model as dev backs sparse MoE: small experts favor inference
lxfater · x · 2026-10-10
Quoting Qwen's cryptic "big or small?" teaser, developer lxfater says he would pick a sparse model, since small-expert MoE architectures are friendlier for inference — only a subset of parameters activates at inference time, cutting serving cost.
More from Models
- Opus 5.5 Fast beats Sol 6.1 Ultrafast on price: 33% cheaper at the same ~300 tok/s speed — kimmonismus · 2026-10-10
- Developer slams Grok's 'nails on chalkboard' chat style and AI slop speech — max_paperclips · 2026-10-10
- Unannounced gpt-rosalind-discovery model spotted on OpenAI's API pricing page at $5/$25 per 1M tokens — testingcatalog · 2026-10-10
- Kimi K3.1 reportedly launching mid-October with long-horizon agent focus; K2.8 pricing raises concerns — teortaxesTex · 2026-10-10
- Kimi K3.1 not launching before mid-October, clarifies ChrisGPT — ChrisGPT · 2026-10-10
- Codex quality dip reported; GPT-61 Sol High said to be the better tier — vista8 · 2026-10-10