Kimi K3 reportedly keeps block attention residuals in a 93-layer design
burny_tech · x · 2026-07-28
A repost highlights a video on attention residuals and Kimi K3’s training/infrastructure details.
- Kimi K3 keeps the block attention residuals design unchanged.
- The model uses 8 blocks with 12 layers each, plus a partial final block, or 9 total blocks if the embedding layer is counted.
- It is reported to have 93 total layers and hidden size 7168.
- The speaker argues the model’s aspect ratio of 77 is conservative and reflects a preference for a deeper, not wider architecture, since attention residuals work better in deeper models.
More from Infra
- Kimi K3 runs on AMD MI350X with SGLang and hits 327 tok/s across four requests — burny_tech · 2026-07-28
- llama.cpp benchmarks show ROCm and Vulkan trading wins on AMD Radeon AI PRO R9700 — Gesha24 · 2026-07-28
- NASA’s new administrator says SpaceX-backed orbital data centers will happen — elonmusk · 2026-07-28
- Reddit users test whether unlocked CMP 170HX cards can power local AI rigs — Tritheone69 · 2026-07-28
- Open-weight models are pitched as a security win for large companies — jessi_cata · 2026-07-28
- Dolphin is being trained on Trinity Large Thinking 398B across 72 RTX 4090s — QuixiAI · 2026-07-28