Gemma 4 is benchmarked locally on a 48GB Mac with MLX, llama.cpp and Java 25
rseroter · x · 2026-07-28
A benchmark post compares three ways to run Gemma 4 26B-A4B locally on an Apple Silicon Mac with 48GB of unified memory.
- The setup targets a fully local workflow from Java 25 via LangChain4j.
- It compares MLX, llama.cpp with MTP speculative decoding, and a pure Java 25 path.
- The post is framed as a practical local-inference benchmark rather than a model announcement.
The quoted thread also notes that the Gemma family has surpassed 900 million downloads.
More from Infra
- fmgo: call Apple's on-device Foundation Models from Go with no CGO and no Swift — Super_Run_8466 · 2026-09-23
- Huawei unveils Peerium architecture: nested BSP unifies million processors into one computer — Dr_Singularity · 2026-09-23
- Grok explains why DeepSeek picked DualPipe + ZeRO-1 over ZeRO-3 on 2048 H800s — TheZachMueller · 2026-09-23
- AI costs fall 47% per quarter, 4x faster than DNA sequencing: Epoch AI — daveholtz · 2026-09-23
- M5 Ultra LLM test: 4x faster prompt processing, but double the power draw — DigitalguyCH · 2026-09-23
- $500 of Dell OptiPlexes become a diskless netboot lab where AI agents can't brick the hardware — colinmcnamara · 2026-09-23