Helion Kernels Land in Hugging Face Kernels: 1.2x Faster Than FlashAttention on H100
RisingSayak · x · 2026-09-15
Helion, Meta's high-level "PyTorch with tiles" DSL for writing performant ML kernels, is now supported in Hugging Face's Kernels project. A joint HF/Meta blog walks through building, autotuning, and shipping pre-tuned kernel configs to the Hub to cut cold-start times. Benchmark: a Helion attention kernel averages 1.2x faster than PyTorch's FlashAttention across 19 tuned shapes on H100.
More from Infra
- OpenAI Used Its Own LLMs to Design Its Jalapeño Chip, Closing an AI-Designs-Chips Loop — satnam6502 · 2026-09-15
- Microsoft reportedly seeks ~38GW of data center capacity by 2032 — VraserX · 2026-09-15
- Memphis residents say noise from Musk data-center power plant is costing them sleep — LuizaJarovsky · 2026-09-15
- SSD prices are rising as AI demand breaks decades of falling storage costs — lemire · 2026-09-15
- Analyst thesis: 2027 is peak chip constraint year, relief not until 2028 — BenBajarin · 2026-09-15
- Full-text search over all 49M HN items runs surprisingly fast on pure Neon Postgres — matei_zaharia · 2026-09-15