Mac mini tested: local 35B runtime hits Haiku-level scores but falls short for agents

PawelHuryn · x · 2026-09-10

This is a runtime, not new models: the 35B is Qwen3.5-MoE 35B-A3B at int4, the 8B is Ling 3.0 tiny. MoE only — pointing it at your own model requires training a prerouter and LoRA adapters first.

Measured on a Mac mini M4 Pro: the 8B runs at 24 tok/s from 4 GB disk / 1 GB RAM (69.9 avg); the 35B at 15 tok/s from 23 GB disk / 2.9 GB RAM (79.2 avg), scoring MMLU-Pro 81.0 and GPQA 79.8 — Haiku territory, well under Sonnet 5.

Not for agents: it's "not yet optimized for agentic tasks," a 3.3k prompt takes 30s to first token, and long contexts blow past the 3 GB memory claim. Good for private one-shot local work on small-RAM Macs.

Original post →

More from Infra

Infra channel →