llama.cpp Adds Maple 20B-A1B Ternary MoE Architecture for CPU and Low-VRAM Devices
jacek2023 · reddit · 2026-09-14
A pull request (#27000) adds support for DeepGrove Maple 20B-A1B's ternary MoE architecture to llama.cpp. The model has 20B total parameters with roughly 1B active, targeting CPU and low-VRAM local deployment; a preview is live on Hugging Face.
More from Infra
- EU RAM prices cool off: DDR5 kits down 6-7% in 30 days, Italy beats Germany — egudegi · 2026-09-14
- Apple's A20 Pro hits ~4000 in Geekbench single-core, 25% ahead of best x86 chips — lemire · 2026-09-14
- Portugal could lift GDP 1-2% by becoming the fastest place to build data centers — dscape · 2026-09-14
- Claude usage boosts quietly removed, fueling talk that 'The Great Compute Crunch has begun' — jacob_posel · 2026-09-14
- Unions urged to halt AI datacenter buildout until jobs and grid use are protected — nordicinst · 2026-09-14
- SK hynix completes HBM4 internal qualification, ushering in custom base die competition — blaizedsouza · 2026-09-14