Maple-Preview: Open-Source Ternary-Weight LLM Hits 200+ tokens/s on Mac Mini
garrytan · x · 2026-08-05
DeepGrove has introduced Maple-Preview, an open-source 20B parameter (1B active) ternary-weight reasoning LLM.
The model achieves state-of-the-art performance within its weight class and is capable of solving IMO-level math problems. Thanks to its ternary weights, it runs at over 200 tokens per second on a Mac Mini M4, making it 5–16× faster than comparable efficient models like Gemma and Qwen.
Related event: DeepGrove's open-source Maple model runs at 200+ tokens/s on Mac Mini(4 posts)→
More from Infra
- Update CUDA to 13.3 to Fix DeepSeek V4 Flash Looping Issue — Easy_Werewolf7903 · 2026-08-05
- Hands-On: Deploying 550B Nemotron 3 Ultra Locally on NVIDIA DGX Station — NVIDIA Developer · 2026-08-05
- Together AI's Monthly Token Volume Skyrockets from 30B to 400T — togethercompute · 2026-08-05
- API Key Expiry Leads to Runaway Agent, Costs $300 in Idle Compute — voooooogel · 2026-08-05
- Dev Builds Pixel-Art GPU Cluster Dashboard in 20 Mins Using GLM Agent — Porespellar · 2026-08-05
- DeepSeek-V4-Flash Runs with 256k Context on 4x 4090 GPUs — dangerous_inference · 2026-08-05