Hugging Face deep dive: browser-compiled GPU kernels, attention in 20 lines of JS at 400 fps

nicodotdev · x · 2026-09-04

Two days after shipping @huggingface/kernels, Hugging Face released a deep dive on the design: they publish Jinja templates instead of WGSL files so the browser compiles the fastest kernel for your GPU. Highlights: attention written in 20 lines of JS running on GPU, a single matmul animating 1M+ pixels at 400 fps vs 6 fps in plain JS, and Fleet, a tool to benchmark your GPU and help tune kernels for every device.

Related event: HF Deep-Dives Into kernels Library That Compiles GPU Kernels in the Browser(2 posts)→

Original post →

More from Infra

Infra channel →