Transformers Now Calls GGML Kernels Directly, Matching llama.cpp Performance for GGUF

On September 22, Lysandre, a core maintainer of Hugging Face Transformers, announced a major upgrade to GGUF support: Transformers can now call GGML kernels directly at runtime via the kernels library, achieving inference performance on par with llama.cpp. Previously, while GGUF files could be loaded, they had to be dequantized first, resulting in noticeable performance losses.

Confirmed

Why it matters

2026-09-22 ~ 2026-09-22 · 6 related posts

Primary sources

2 near-duplicate retellings: LysandreJik · LysandreJik