Gemma 4 is benchmarked locally on a 48GB Mac with MLX, llama.cpp and Java 25
rseroter · x · 2026-07-28
A benchmark post compares three ways to run Gemma 4 26B-A4B locally on an Apple Silicon Mac with 48GB of unified memory.
- The setup targets a fully local workflow from Java 25 via LangChain4j.
- It compares MLX, llama.cpp with MTP speculative decoding, and a pure Java 25 path.
- The post is framed as a practical local-inference benchmark rather than a model announcement.
The quoted thread also notes that the Gemma family has surpassed 900 million downloads.
More from Infra
- B200 fine-tuning benchmark shows up to 5.8× speedup and 50% less VRAM — antoine_chaffin · 2026-07-28
- Packed-encoders claims 5× faster training by eliminating padding FLOPs — antoine_chaffin · 2026-07-28
- Bittensor subnet expansion is pitched as a cheaper AI infrastructure path for companies — markjeffrey · 2026-07-28
- Project Orion is training a 16B model live across three continents on heterogeneous compute — markjeffrey · 2026-07-28
- Modular handbook maps the hidden costs of LLM inference, from KV cache to prefill/decode splits — udmrzn · 2026-07-28
- Kimi K3 reaches Merge Gateway with U.S. inference providers and ZDR terms — shensi · 2026-07-28