llama.cpp Metal optimization boosts IQ3_XXS decode speed on Apple Silicon
predatar · reddit · 2026-09-01
A submitted PR optimizes the Metal backend in llama.cpp, improving decode speed for IQ3XXS models on Apple Silicon. Benchmarks show TieL Coder 35B A3B decode speed increased from 65.6 to 73.9 tok/s. The author invites community testing and teases upcoming prefill improvements.
More from Infra
- Nvidia Earnings: Avoiding Consolidation and Dollars per Gigawatt — Stratechery · 2026-09-01
- TEAS benchmark: measures inference at natural lengths, reporting cost, accuracy, and energy — PontiEdoardo · 2026-09-01
- Nebius GTM on AI Bottlenecks and Inference Demand Explosion — demian_ai · 2026-09-01
- Beyond Model Speed: 19 Distributed Patterns to Optimize AI Latency — bibryam · 2026-09-01
- MongoDB CTO on Database Architecture Evolution and the Unsolved Problem of Agent Memory — The Cognitive Revolution · 2026-09-01
- MongoDB's Pete Johnson on How Retrieval Drives Agent Performance — The Cognitive Revolution · 2026-09-01