llama.cpp merges new Metal kernels, making speculative decoding 3.4x faster on M3 Ultra

ggerganov · x · 2026-10-05

Original post →

More from coding & agent

coding & agent channel →