Agentic kernel optimization lifts a Kimi-like model from 65 to 406 tok/s

xenovatech · reddit · 2026-07-25

The post shows an “agentic kernel optimization” run where a swarm of GPT-5.6 Sol agents spent more than 40 hours improving a Kimi K3-like model from 65 tok/s to 406 tok/s.

The animation tracks how they fused operators, transformed the execution graph, and reduced a fully decomposed 331-node Kimi Linear graph to one that needs only 22 GPU dispatches per token. The biggest gains came from custom WebGPU kernels for KDA, MLA, and MoE, with a full framework release promised later.

Original post →

More from Infra

Infra channel →