Launch HN: Magnitude (YC S25) ships self-optimizing inference engine, up to 2x faster than llama.cpp
anerli · hn · 2026-10-01
Magnitude (YC S25), from two engineers who previously built an open-source browser agent with 4k+ GitHub stars, launches a self-optimizing inference engine for local agents, claiming up to 2x speedups over llama.cpp on any Mac/Linux/Windows hardware.
- The gap: existing engines trade off—vLLM/SGLang optimize batched datacenter inference at the cost of single-session speed; llama.cpp/Ollama prioritize compatibility over hardware-specific performance; specialized engines lack completeness. None is designed for agent workloads (long, concurrent sessions that must coexist with normal desktop use).
- Approach: kernels are written with tunable parameters and compiled/tuned on the user's actual device, combining broad compatibility with hardware-specific performance ceilings; efficient kernels target only the most popular open-weight model families; dynamic memory allocation reserves only weight memory up front and grows/frees the heap as agent sessions start and end.
More from coding & agent
- SkillSeek: plain BM25 matches LLM-mediated agent skill retrieval at half the cost — StevensAGI · 2026-10-01
- Recreating all five Dot characters in real time with SDF primitives — yihui_indie · 2026-10-01
- Editor open-sources open-fusion-mcp: Claude builds editable motion graphics inside DaVinci Resolve — JohnnyLegion · 2026-10-01
- codemode + general classification models demoed in pi draws developer praise — ricklamers · 2026-10-01
- Ex-Cursor engineer runs 6 Grok bots: from prompting to hiring a bot team — lasas · 2026-10-01
- Agent-built custom Lego sets: dev lets AI design and order real sets — noahsolomon · 2026-10-01