Moonshot open-sources MoBA, a block attention mechanism it says is 16x faster for long context
gekobraa · x · 2026-09-09
Moonshot AI (Kimi's maker) has open-sourced MoBA (Mixture of Block Attention), the attention mechanism behind Kimi's long context, claiming 95% of standard attention computation is wasted.
- Vanilla attention scales quadratically ($O(N^2)$): every token attends to all previous tokens
- MoBA applies Mixture-of-Experts ideas to the attention layer: context is chopped into blocks, and a lightweight dynamic gating lets each query token pick relevant blocks instead of brute-forcing full attention
- Reported up to 16x faster than full attention
More from Models
- Claims surface that Hugging Face was hacked two months ago via decades-old Linux bug — Justin_Halford_ · 2026-09-09
- OpenAI may pause new Pro subscriptions if GPT-6 usage keeps growing, says Tibo — op7418 · 2026-09-09
- OpenAI may pause new Pro subscriptions as demand for Astra hits unprecedented levels — op7418 · 2026-09-09
- Veteran claims a trained intuition for sniffing out AI-written text — IndraVahan · 2026-09-09
- Tencent Hunyuan details Gander, an end-to-end full-duplex omni interaction agent — Tencent-Hunyuan · 2026-09-09
- From Failing 2+2 to PhD-Level Math: A Timeline of AI's Three-Year Leap — haider1 · 2026-09-09