20B Model Maple-Preview Runs at 200+ tokens/s on Mac Mini, Solves IMO Math
tylerbruno05 · x · 2026-08-05
DeepGrove has introduced Maple-Preview, an open-source 20B-A1B ternary-weight reasoning LLM, claiming SOTA performance in its weight class.
The model showcases extreme edge-inference efficiency:
- High Speed: Achieves 200+ tokens/sec on a Mac Mini M4, running 5–16× faster than efficient models like Gemma 4 and Qwen3.5.
- Low Footprint: Requires only 5.3 GB of memory.
- Strong Capabilities: Capable of solving IMO-level math problems.
- On-Device Training: Can literally train itself on-device to adapt to user preferences.
More from Infra
- Elon Musk Expects First Starmind AI Satellites to Launch Next Year — Polymarket · 2026-08-05
- GPU Rental Costs Surge: Blackwell Clusters Up 60% in Six Months — Alec_Coughlin · 2026-08-05
- Musk: SpaceX to Triple GPU Compute by 2027, Going All-In on Nvidia — firstadopter · 2026-08-05
- Elon Musk says SpaceX committed to using Nvidia GPUs exclusively — elonmusk · 2026-08-05
- 25% Chance of an AI Data Center in Space by 2027, Polymarket Odds Show — Polymarket · 2026-08-05
- Report: SpaceX to Use Nvidia Chips for Orbital AI Data Center Satellites — Polymarket · 2026-08-05