Transformers Now Runs GGUF at llama.cpp Speed via GGML Kernels
LysandreJik · x · 2026-09-22
Hugging Face maintainer Lysandre announced that Transformers, which previously supported GGUF files only by unquantizing them, can now run them with GGML kernels through the kernels library, matching llama.cpp performance. The work was contributed by @marcsun, with kernels provided by the ggmlorg team.
More from Infra
- Local H3 video gen: 15s clip in ~5 min at zero cost vs ~$0.60 per generation on platforms — AIandDesign · 2026-09-22
- MotherDuck makes text classification 50x faster at ~1% of LLM cost — josh_wills · 2026-09-22
- Google serves 3.2 quadrillion AI tokens a month — ~12GW, and energy could run out in 3 years — victor_explore · 2026-09-22
- M5 Ultra 256GB early test: Mimo2.6-Flash hits 49 tok/s at 32K context — bakawolf123 · 2026-09-22
- TRL async GRPO adds LoRA sync via storage bucket and proxy, cutting 500-step training to 53 min — SergioPaniego · 2026-09-22
- GGUF models can now run directly in Hugging Face transformers with ggml Metal kernels — pcuenq · 2026-09-22