Hugging Face Transformers now runs llama.cpp GGUF quants natively on your laptop
ariG23498 · x · 2026-09-29
Hugging Face announced native GGUF support in transformers: pick any GGUF checkpoint from the Hub and load it with frompretrained, running quants sized for your laptop through the familiar transformers API. Key points:
- Reuses llama.cpp's underlying ggml kernels (via the kernels library) and cuts overhead in generate to approach llama.cpp-level performance
- Initial focus is local inference on Apple Silicon
- Unsloth, LM Studio Community and bartowski already publish GGUF checkpoints in many quantization levels
- transformers serve can plug into local tools like Pi or Jan
- Cites Julien Chaumond's demo: Qwen3.6 27B via llama.cpp running the Pi coding agent on a MacBook Pro, feeling close to the latest Opus on non-trivial HF codebase tasks
A significant ecosystem upgrade for local deployment and quant debugging/eval.
More from Infra
- antirez: give companies a fleet of DGX Sparks and Mac Ultras to survive API quota limits — antirez · 2026-09-29
- Ornith-1.5 releases DFlash draft checkpoints for 9B/397B/35B-A3B speculative decoding — jacek2023 · 2026-09-29
- Jensen Huang says Nvidia stock is undervalued, plans hundreds of billions in buybacks — firstadopter · 2026-09-29
- Uber Eats ranking models serve 8M predictions/sec: how Uber scales ML feature consistency — AxSaucedo · 2026-09-29
- Qdrant unveils Constella research preview: swap query embedding models without re-embedding your docs — qdrant_engine · 2026-09-29
- 124M model with a 65B embedding sparks the AFED disaggregation joke — YouJiacheng · 2026-09-29