128GB Mac Studio M5 Max tested: Gemma 4 26B-A4B hits 134 tok/s and aces all 12 tasks
TheOyinbooke · x · 2026-10-03
The author benchmarked seven models that never fit on a 16GB machine, plus a Gemma 3 4B baseline, on a Mac Studio M5 Max with 128GB unified memory and a 40-core GPU — measuring tokens/sec, memory, and strict pass/fail correctness (chat, summary, code against hidden tests, one reasoning question), median of three runs each.
Key findings:
- Gemma 4 26B-A4B is the clear pick: 134 tok/s on short answers, 123 on long summaries, 12/12 correct in every run, at just 15.4GB. The only model both fast and always right.
- Two Splash-engine models were faster (one hit 361 tok/s) but required installing a second LM Studio app and scored lower on correctness.
- "A fast wrong answer is still a wrong answer" — all eight models (222GB total) ran from an external SSD.
More from Infra
- AI Conference Day 2: Agent Inference Costs and Observability Steal the Show — sanjaykalra · 2026-10-03
- Apple M5 Ultra 256GB Context Benchmark Shows Strong Prefill Speed on Quantized Qwen — antirez · 2026-10-03
- Fractile Founder on Winning the AI Chip Market and the Memory Bandwidth Bottleneck — saranormous · 2026-10-03
- SoftBank pays $3.1B for data-center investment manager as buildout runs on debt — YvesMulkers · 2026-10-03
- Cloudflare Open-Sources Its Internal AI 'OS', Hits 10k GitHub Stars — ritakozlov · 2026-10-03
- CoreWeave wants the industry to stop calling it a 'neocloud': 'you can only be new for a short time' — DavidLinthicum · 2026-10-03