Magnitude (YC S25) Open-Sources Inference Engine with Up to 2x Faster Decode Than llama.cpp
petrusenko_max · x · 2026-10-01
Magnitude (YC S25) released an open-source inference engine that compiles and tunes kernels on-device, reporting up to 2x faster decode than llama.cpp: +92% on Metal, +19% on CUDA. It runs on Apple Silicon, NVIDIA, AMD, or CPU-only hardware under Apache-2.0.
More from Infra
- Swarms Rust claims 130-440x faster startup than LangChain, LangGraph and CrewAI — KyeGomezB · 2026-10-01
- India's EtherealMachine Builds Its Own 5-Axis CNC Machines From Scratch — RoboBalaji · 2026-10-01
- Merge Agent Handler ships as a partner recipe in NVIDIA NemoClaw for safe enterprise agents — shensi · 2026-10-01
- Google's Data Agent Kit hits GA, wiring 15+ data services into your coding agent via MCP — rseroter · 2026-10-01
- The agent loop is what matters: why local LLMs keep you in control — ag789 · 2026-10-01
- Framework opens preorders for AMD Ryzen AI Max 400 desktop with 192GB RAM — Educational_Sun_8813 · 2026-10-01