Hugging Face deep dive: browser-compiled GPU kernels, attention in 20 lines of JS at 400 fps
nicodotdev · x · 2026-09-04
Two days after shipping @huggingface/kernels, Hugging Face released a deep dive on the design: they publish Jinja templates instead of WGSL files so the browser compiles the fastest kernel for your GPU. Highlights: attention written in 20 lines of JS running on GPU, a single matmul animating 1M+ pixels at 400 fps vs 6 fps in plain JS, and Fleet, a tool to benchmark your GPU and help tune kernels for every device.
Related event: HF Deep-Dives Into kernels Library That Compiles GPU Kernels in the Browser(2 posts)→
More from Infra
- Perplexity CEO pitches local AI hardware with Nvidia and Apple, backs hybrid cloud-local — Kr00ney · 2026-09-05
- Apple vs. NVIDIA: the balance sheet as competitive advantage, B2C then, B2B2B now — BenBajarin · 2026-09-05
- Anima's Accelerated Understanding Bets on a Foundation Model for the Physical World — Latent Space · 2026-09-05
- Dev forks vLLM with custom patch to benchmark 31B model unsupported by flashinfer — abhijithneil · 2026-09-04
- Investor Gavin Baker: AI data centers are 'the best thing' for US working class — GavinSBaker · 2026-09-04
- Kafka, Kafka Connect and Schema Registry exposed as native MCP tools — jkriket · 2026-09-04