LLM Architecture: How Large Vocabularies Mitigate Softmax Rank Bottlenecks
kalomaze · x · 2026-08-14
Researcher @kalomaze discussed the Softmax rank bottleneck issue in large language model architectures like DeepSeek V3.
- Context: In a reply, it was noted that models seem heavily designed to avoid the softmax rank bottleneck, otherwise some nats can objectively never be modeled correctly.
- Analysis: He pointed out that DSV3-like backbones have hidden dimensions of around 7168. With a 15k vocabulary size, the softmax is 10x less rank-bottlenecked in terms of the finegrainedness of the learned probability distribution compared to narrower architectures.
More from Research
- Researcher Rants: The Term 'Emergence' Masks Our Lack of Causal Understanding — BasedRaddka · 2026-08-14
- WARP: Transforming Offline Human Motion into Replayable Whole-Body Robot Actions — chris_j_paxton · 2026-08-14
- Yi Ma to Keynote Workshop on Mathematical Foundations of AI at NUS — YiMaTweets · 2026-08-14
- New Paper Explains How LLMs 'Hijack' Pleistocene Brains into Perceiving False Agency — MacrinePhD · 2026-08-14
- Nanjing University's Marope Framework Enables Robots to Jump Rope with Humans — ericjang11 · 2026-08-14
- REKEY: New Benchmark Exposes VLM Score Inflation from Memorization — jiqizhixin · 2026-08-14