llama.cpp Adds Maple 20B-A1B Ternary MoE Architecture for CPU and Low-VRAM Devices

jacek2023 · reddit · 2026-09-14

A pull request (#27000) adds support for DeepGrove Maple 20B-A1B's ternary MoE architecture to llama.cpp. The model has 20B total parameters with roughly 1B active, targeting CPU and low-VRAM local deployment; a preview is live on Hugging Face.

Original post →

More from Infra

Infra channel →