Kimi Ditches Positional Encoding Entirely at Scale
tokenbender · x · 2026-07-28
Developers analyzing Kimi's large model architecture discovered that it completely removes Rotary Position Embedding (RoPE) at its current scale, eliminating this baked-in positional bias. The author views this as another victory for the 'Bitter Lesson' in AI scaling.
Related event: Kimi K3 Architecture Drops RoPE for NoPE(3 posts)→
More from Models
- Bug Hunt Bench ranks GPT-6 Astra top as coding models fix real planted bugs, costs spread 200x — PawelHuryn · 2026-09-23
- Third-party test: Claude Opus 5.5 renders finer 3D scenes but costs 13x more than GPT-6 Sol — testingcatalog · 2026-09-23
- Tester claims Claude Opus 5.5 has the best visual design output of any model tested — burny_tech · 2026-09-23
- GPT-6 Sol Codex system prompt leaked: over 294,000 characters dumped on GitHub — gaganghotra_ · 2026-09-23
- Claude 5.5 (live) keeps generating user turns, reports user — BlackHC · 2026-09-23
- Code benchmarks are mostly slop: dev calls for narrow evals per domain, not one score — almmaasoglu · 2026-09-23