Magnitude (YC S25) Open-Sources Inference Engine with Up to 2x Faster Decode Than llama.cpp

petrusenko_max · x · 2026-10-01

Magnitude (YC S25) released an open-source inference engine that compiles and tunes kernels on-device, reporting up to 2x faster decode than llama.cpp: +92% on Metal, +19% on CUDA. It runs on Apple Silicon, NVIDIA, AMD, or CPU-only hardware under Apache-2.0.

Original post →

More from Infra

Infra channel →