Transformers Now Calls GGML Kernels Directly, Matching llama.cpp Performance

LysandreJik · x · 2026-09-22

Hugging Face's Lysandre announced Transformers can now use GGML kernels via the kernels library at the same performance as llama.cpp — previously it only loaded GGUF files by unquantizing them. As proof of concept, only Qwen 3.5 and its MoE counterpart are supported; users can request models via GitHub issues.

More exciting: models outside llama.cpp — even without GGUF files — can leverage GGML kernels. LLM coverage in llama.cpp is already huge, but modalities with heavy pre/post-processing (raw CV, diffusion, TTS) could benefit most. Still PoC, but promising.

Related event: Transformers Now Calls GGML Kernels Directly, Matching llama.cpp Performance for GGUF(6 posts)→

Original post →

More from Infra

Infra channel →