Transformers Now Runs GGUF at llama.cpp Speed via GGML Kernels

LysandreJik · x · 2026-09-22

Hugging Face maintainer Lysandre announced that Transformers, which previously supported GGUF files only by unquantizing them, can now run them with GGML kernels through the kernels library, matching llama.cpp performance. The work was contributed by @marcsun, with kernels provided by the ggmlorg team.

Related event: Transformers Now Calls GGML Kernels Directly, Matching llama.cpp Performance for GGUF(6 posts)→

Original post →

More from Infra

Infra channel →