GGUF models can now run directly in Hugging Face transformers with ggml Metal kernels

pcuenq · x · 2026-09-22

GGUF checkpoints from llama.cpp can now run natively in Hugging Face transformers. The work ports ggml's Metal kernels into the transformers ecosystem, improving compatibility and enabling faster local inference on Mac — millions of downloaded llama.cpp models now have one more way to run without conversion.

Related event: Transformers Now Calls GGML Kernels Directly, Matching llama.cpp Performance for GGUF(6 posts)→

Original post →

More from Infra

Infra channel →