GLM-4.5-Air now supports MTP acceleration in llama.cpp
jacek2023 · reddit · 2026-08-24
A user highlights that the older GLM-4.5-Air model can now achieve significant speedups by enabling MTP (Multi-Token Prediction) in llama.cpp. It is a 106B MoE model with only 12B active parameters, making it suitable for hardware with high memory but limited compute (e.g., RTX 3090). The user provides a link to the MTP GGUF file and recommends several creative writing/RP finetunes available on Hugging Face.
More from Models
- Anthropic Addresses Opus Verbosity with Config Command — kimmonismus · 2026-08-24
- Uncensored Qwen3.8-27B Model Trends on Hugging Face — orcarouter · 2026-08-24
- SemiAnalysis: $200/mo AI plans can yield up to $8K-$14K in tokens — scaling01 · 2026-08-24
- Voice AI Startup Modulate Takes #1 Spot on Hugging Face's Transcription Benchmark — jonathan_wilke · 2026-08-24
- Ox Alpha questioned as overhyped; early tests show non-frontier performance — peterwildeford · 2026-08-24
- Rumor: Anthropic testing Haiku 5 and Sonnet 5.1 instead of Fable 5.1 — ChrisGPT · 2026-08-24