Mac mini tested: local 35B runtime hits Haiku-level scores but falls short for agents
PawelHuryn · x · 2026-09-10
This is a runtime, not new models: the 35B is Qwen3.5-MoE 35B-A3B at int4, the 8B is Ling 3.0 tiny. MoE only — pointing it at your own model requires training a prerouter and LoRA adapters first.
Measured on a Mac mini M4 Pro: the 8B runs at 24 tok/s from 4 GB disk / 1 GB RAM (69.9 avg); the 35B at 15 tok/s from 23 GB disk / 2.9 GB RAM (79.2 avg), scoring MMLU-Pro 81.0 and GPQA 79.8 — Haiku territory, well under Sonnet 5.
Not for agents: it's "not yet optimized for agentic tasks," a 3.3k prompt takes 30s to first token, and long contexts blow past the 3 GB memory claim. Good for private one-shot local work on small-RAM Macs.
More from Infra
- Matt Barrie burned 4B tokens in a day, cut his bill 500-fold, and now worries about $5T in debt — gaganghotra_ · 2026-09-10
- Analyst: DeepSeek's latest change is a big win for token efficiency, moving toward OpenAI's regime — teortaxesTex · 2026-09-10
- Acellera tests 7 LLM+harness combos on drug discovery: one RTX 5090 holds up — gdefabritiis · 2026-09-10
- Miles ships Day-0 RL support for DeepSeek-V4.1-Flash with KL held at 0.0012–0.0017 — ying11231 · 2026-09-10
- The data center is a symbol: why debunked claims about AI infrastructure still spread — ShakeelHashim · 2026-09-10
- Kimi K3 lands on RunPod: 2.8T params, 1M context, $3/$15 per 1M tokens — Kimi_Moonshot · 2026-09-10