GGUF models can now run directly in Hugging Face transformers with ggml Metal kernels
pcuenq · x · 2026-09-22
GGUF checkpoints from llama.cpp can now run natively in Hugging Face transformers. The work ports ggml's Metal kernels into the transformers ecosystem, improving compatibility and enabling faster local inference on Mac — millions of downloaded llama.cpp models now have one more way to run without conversion.
More from Infra
- bitsandbytes2 slightly delayed: zero-config lazy compression plus dynamic expert swaps for near-infinite KV cache — Tim_Dettmers · 2026-09-22
- Iterative sensitivity probing: how bitsandbytes2 finds each layer's compression limit — Tim_Dettmers · 2026-09-22
- Tim Dettmers releases runtime dynamic compression framework, hits 1.5-2.0 bit at high quality — Tim_Dettmers · 2026-09-22
- PSA: Non-US Users Should Consider Local AI in Case Governments Ban LLMs — TheMoonMidas · 2026-09-22
- Engram: A Local Encrypted Memory Vault Unifying Agent Memory Across AI Tools — Acceptable_Leg3950 · 2026-09-22
- DigitalOcean Managed Agents enters public preview with idle-pause billing — damianplayer · 2026-09-22