Agentic kernel optimization lifts a Kimi-like model from 65 to 406 tok/s
xenovatech · reddit · 2026-07-25
The post shows an “agentic kernel optimization” run where a swarm of GPT-5.6 Sol agents spent more than 40 hours improving a Kimi K3-like model from 65 tok/s to 406 tok/s.
The animation tracks how they fused operators, transformed the execution graph, and reduced a fully decomposed 331-node Kimi Linear graph to one that needs only 22 GPU dispatches per token. The biggest gains came from custom WebGPU kernels for KDA, MLA, and MoE, with a full framework release promised later.
More from Infra
- A shadow server can virtual-patch legacy power systems that cannot be fixed directly — lauriewired · 2026-07-25
- Cloudflare launches x402-powered payments for AI agents on July 1 — kleffew94 · 2026-07-25
- Oil services giant SLB is betting on AI data centers and a rebound in energy demand — fortune · 2026-07-25
- PyTorch Foundation expands its neutral AI stack to vLLM, DeepSpeed and Ray — PyTorch · 2026-07-25
- Lightning AI says its interconnect is now faster than bare metal — LightningAI · 2026-07-25
- Lightning Cloud adds S3 connectivity and faster multi-node training — LightningAI · 2026-07-25