GLM-4.5-Air now supports MTP acceleration in llama.cpp

jacek2023 · reddit · 2026-08-24

A user highlights that the older GLM-4.5-Air model can now achieve significant speedups by enabling MTP (Multi-Token Prediction) in llama.cpp. It is a 106B MoE model with only 12B active parameters, making it suitable for hardware with high memory but limited compute (e.g., RTX 3090). The user provides a link to the MTP GGUF file and recommends several creative writing/RP finetunes available on Hugging Face.

Original post →

More from Models

Models channel →