Hugging Face Kernels quickstart: load GPU-optimized kernels in one line
ariG23498 · x · 2026-10-06
Dev ariG23498 kicks off a series on compute kernels — instruction lists that run in parallel across many GPU compute units (e.g. NVIDIA SMs). The linked Hugging Face Kernels quickstart shows how to fetch optimized kernels from the Hub with getkernel("kernels-community/activation", version=1) and run ops like gelufast on CUDA in a few lines. Kernels are versioned by major version; within a branch the API never breaks and older PyTorch builds stay available — community-optimized kernels without writing CUDA.
More from coding & agent
- Developer moves all work to Amp Code orbs, asks how to manage env secrets — iannuttall · 2026-10-06
- Vite's bundled dev, one flag, no config: 21x fewer requests, 2-4x faster HMR on tldraw — cnakazawa · 2026-10-06
- Matt Pocock shares prompt that uses 3 subagents to radically restructure your AGENTS.md — mattpocockuk · 2026-10-06
- Matt Pocock launches The AI Coding Dictionary to standardize AI coding terms like harness, spec and cache tokens — mattpocockuk · 2026-10-06
- These agents ate 780GB of disk space in 3 weeks — kevinkern · 2026-10-06
- Google's AIM paper: research agents improve faster by mapping and auditing ideas, beating baselines up to 3.1x sooner — rohanpaul_ai · 2026-10-06