Helion Kernels Land in Hugging Face Kernels: 1.2x Faster Than FlashAttention on H100

RisingSayak · x · 2026-09-15

Helion, Meta's high-level "PyTorch with tiles" DSL for writing performant ML kernels, is now supported in Hugging Face's Kernels project. A joint HF/Meta blog walks through building, autotuning, and shipping pre-tuned kernel configs to the Hub to cut cold-start times. Benchmark: a Helion attention kernel averages 1.2x faster than PyTorch's FlashAttention across 19 tuned shapes on H100.

Original post →

More from Infra

Infra channel →