Hugging Face Transformers now runs llama.cpp GGUF quants natively

ariG23498 · x · 2026-09-23

Hugging Face announces that transformers now natively supports GGUF models: pick a GGUF from the Hub and load it with the familiar frompretrained API to generate locally.

The article also cites Julien Chaumond's demo of Qwen3.6 27B running in the Pi coding agent via llama.cpp on a MacBook Pro, performing close to the latest Opus.

Related event: Hugging Face Transformers Now Runs GGUF at llama.cpp Speed via GGML Kernels(8 posts)→

Original post →

More from Infra

Infra channel →