Alibaba's Justin Lin: huge sparse models take up memory but won't fix reasoning
JustinLin610 · x · 2026-10-03
Qwen lead Justin Lin (JustinLin610) said on X he has no current interest in a large, sparse model despite needing to learn more about it. His reasoning: it is "too big and too sparse, taking up memory and storage," and while it may benefit people, it likely won't help reasoning — which he sees as today's bottleneck for intelligence.
More from Models
- GPT learns to pass tests, not to engineer: the RL reward-mismatch behind ugly code — xiaohu · 2026-10-03
- Local LLM setup: Strata lets a 7900XTX + 64GB RAM run Qwen Flash at 60 tok/s at 250K context — soyalemujica · 2026-10-03
- User reports OpenAI API account with $2,000+ spend wiped clean, Codex credits zeroed — RecursivelyYours · 2026-10-03
- Yoav Goldberg: the 'post-training adds no knowledge' dogma is clearly no longer true — yoavgo · 2026-10-03
- CASIA open-sources 9B ZDTaichu5.0, a multimodal model for 3D spatial and embodied reasoning — mikelau2026 · 2026-10-03
- Ant's InclusionAI ships Ling-3.1-flash, a ~560B MoE agent model with open weights promised — mikelau2026 · 2026-10-03