Kernel Fusion Boosts 975B Parameter Model Speed by 18%

yisongyue · x · 2026-08-12

AsariAILabs shared optimization practices for Inkling, a 975B parameter model developed by thinkymachines. By applying kernel fusion at concurrency c=32, the generated kernels achieved speedups across the entire c=1 to 128 range.

This optimization improved both interactivity and throughput, with performance gains reaching up to 18%.

Original post →

More from Infra

Infra channel →