Kimi Ditches Positional Encoding Entirely at Scale
tokenbender · x · 2026-07-28
Developers analyzing Kimi's large model architecture discovered that it completely removes Rotary Position Embedding (RoPE) at its current scale, eliminating this baked-in positional bias. The author views this as another victory for the 'Bitter Lesson' in AI scaling.
More from Models
- vLLM Collaborates with DigitalOcean to Host Kimi K3 Model — vllm_project · 2026-07-28
- Microsoft says MAI-Cyber-1-Flash hits 96% on CyberGym at half the cost — mustafasuleyman · 2026-07-28
- Kimi K3 lands on ChatLLM with U.S. hosting and an open-source fine-tune — bindureddy · 2026-07-28
- Kimi K3 report introduces SiTU-GLU, a bounded tanh × sigmoid gate that approximates SwiGLU — stochasticchasm · 2026-07-28
- Kimi K3’s MoE routing may be driving higher expert-parallel communication costs — stochasticchasm · 2026-07-28
- Kimi K3 finds 16 new vulnerabilities and beats GLM-5.2 on an exploit benchmark — zephyr_z9 · 2026-07-28